DevOps Engineer
About Re:Build
At Re:Build, our mission is to ensure the next generation of important products are made, at scale, in America. We are laying the foundation for a better future for our customers, employees, and communities by revitalizing America's manufacturing base and creating meaningful jobs across the country, including in historically deindustrialized regions.
We operate an advanced, end-to-end manufacturing platform that partners with industrial companies and innovators to take products from first concept to full-scale production in critical verticals including aerospace and defense, electrification, medical, energy and environment, and robotics and automation.
The way we operate is as important as the work we do. It's guided by The Re:Build Way, 16 principles that shape how we collaborate with each other, partner with our customers and vendors, and contribute to the communities where we operate. (link to The Re:Build Way principles)
Who we are looking for
This engineer supports deployments and keeps production running in Cadonix’s controlled AWS environment, including AWS GovCloud (US), where only US Persons are permitted under ITAR. ProduCt teams manage their own continuous integration and delivery workflows and publish artifacts. This role moves those artifacts into the restricted production environment. It verifies them and resolves infrastructure problems from new deployments or spontaneous issues in running systems. Beca or equivalent experience required.use access to this environment is limited, the role is often the only person able to see what is actually happening in production.
What you get to do
- Deployment Support – Assist with and carry out releases of product team materials into the regulated production environment. PPerform pre-deployment verification, run the deployment, check service stability afterward, and complete rollback when a release does not behave as encouraged. Collaborate with product teams to address deployment failures, providing feedback on environmental needs so their pipelines generate artifacts that deploy efficiently.
- AWS infrastructure problem-solving – Investigate and resolve production infrastructure issues across AWS platforms, including compute, container services, networking and security groups, load balancers, DNS and certificates, IAM permissions, storage, and managed databases. Address problems that arise immediately after deployment. Also handle those that occur unexpectedly in a stable system. Determine the root cause instead of just restoring service.
- Production Operations and Incident Management – Coordinate the stability of production systemsCoordinate the management of alerting, dashboards, anHandle incidents, share progress with engineering and product colleagues, and join an on-call rotation.n on-call rotation. Write and maintainDevelop and update runbooks to make recurring issues standard procedure.
- Environment Maintenance – Preserve the environment’s integrity and timeliness: operating system and foundational image updates, certificate renewal, backup and restore verification, capacity headroom, and cost monitoring. AImplement modifications to existing infrastructure using infrastructure as code and version control instead of manual console edits.
- Functioning Inside the Access Boundary – Work in an environment where data remains secure and only authorised US Persons have entry. Handle records, diagnostic information, and support artifacts in line with export-control and customer data-handling requirements. Where a prWhen a problem requires help from an external software vendor, develop sanitized evidence and reproductions. This allows them to assist without accessing controlled or customer data. Raise the issue through accurate channels instead of bypassing boundaries.
- Documentation and Knowledge Exchange – Record environment configuration, deployment procedures, and incident history to ensure operational knowledge is accessible beyond a single individual.Deployment Support – Support and complete deployments of artifacts from the product group into the controlled production. Conduct pre-deployment checks, carry out the deployment, verify service health afterward, and initiate rollback if a release does not behave as encouraged. . Collaborate alongside product teams to resolve deployment failures, feeding back what the environment requires so their pipelines produce artifacts that deploy cleanly.
- AWS Infrastructure Troubleshooting – Diagnose and resolve production infrastructure problems across AWS: compute, container services, network configurations and firewall rules, load balancing mechanisms, DNS and certificates, IAM permissions, storage, and managed databases. Work issues that appear immediately after a deployment as well as those that arise spontaneously in a stable system, and identify root cause rather than only restoring service.
- Production Operations and Incident Response – Monitor production health. Maintain alerting, dashboards, and log collection. Respond to incidents, communicate status to engineering and product stakeholders, and participate in an on-call rotation. Write and maintain runbooks so recurring problems become routine.
- Environment Maintenance – Keep the environment healthy and current: OS and base image patching, certificate updating, backup and restore verification, capacity headroom, and cost monitoring. Apply changes to existing infrastructure through infrastructure as code and version control rather than manual console edits.
- Working Within the Access Boundary – Operate within an environment where data cannot leave and access is limited to individuals authorized under US regulations. Handle logs, diagnostics, and support artifacts consistent with export-control and customer data-handling requirement. When a problem requires help from an external software vendor, build sanitized evidence and reproductions. This lets them assist without accessing controlled or customer data. Bring up the issue using the accurate channels rather than circumventing the boundaries.
- Documentation and Information Sharing – Record system setup, release processes, and incident records to guarantee operational knowledge is available to more than just one person.
What you bring to the Team
- A bachelor’s degree or equivalent experience in Computer Science, Software Engineering, or a related field. Minimum of 5 years’ experience in DevOps, cloud infrastructure or production operations, including at least 3 years supporting production workloads on AWS.
- AWS Fixing – Practical, hands-on experience diagnosing real production problems on AWS: EC2, ECS or EKS, VPC network architecture and security configurations, ALB/NLB, Route 53, ACM, IAM policy evaluation, S3, RDS, and CloudWatch. Capable of reading CloudTrail and CloudWatch logs to reconstruct what happened. AWS GovCloud experience is a plus but not required; equivalent experience in isolated or regulated environments is equally relevant.
- Deployment and Release Mechanics – Understands container images, artifact versioning, environment configuration and secrets injection, and rollback. Comfortable working with CI/CD pipelines owned by other teams and diagnosing where a handoff has gone wrong.
- Infrastructure defined through code – Able to read, modify, and safely apply existing Terraform, CDK, or CloudFormation. Understands state, drift, and why manual console changes cause problems later. Does not need to have designed a large estate from scratch.
- Linux and Networking – Solid Linux administration and fixing. Working understanding of TCP/IP, DNS, TLS, routing, and firewall behaviour — enough to tell a network problem from an application problem.
- Scripting – Proficient in Bash and Python (or Go) for automation, diagnostics, and small tooling.
- Monitoring and Observability – CloudWatch, Prometheus, Grafana, or similar. Able to build a dashboard or alert that answers a real operational question.
- Closed System Problem Solving – Comfortable supporting software they did not write and may not have source access to, working from telemetry, timing, and behaviour rather than code. Persistent and methodical when the obvious answer is not available.
- Data Field – Good judgement about what constitutes sensitive or client-related information, including inside logs, stack traces, configuration, and crash artifacts. Holds the boundary under production pressure and advances rather than improvising. Familiarity with ITAR, CUI, FedRAMP, PCI-DSS, or HIPAA environments is an advantage.
- Communication and Ownership – Keeps collaborators informed during an incident without being asked. Works effectively alongside product groups across time zones, and with external vendors under support contracts. Reliable and self-directed, since much of the environment cannot be observed by anyone else.
The BIG payoff
We are a company who is going to make a difference in the industries and the communities in which we choose to operate. Every employee of Re:Build will share ownership in the company and will share in the financial rewards of the success we achieve together, at all levels of the company!
We want to work with people that reflect the communities in which we operate
Re:Build Manufacturing is proud to be an Equal Employment Opportunity and Affirmative Action employer. We do not discriminate based upon race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, marital status, parental status, cultural background, organizational level, work styles, tenure and life experiences. Or for any other reason.
Re:Build is committed to providing reasonable accommodations for qualified individuals with disabilities in our job application procedures. If you need assistance or an accommodation due to a disability, you may contact us at accommodations.ta@ReBuildmanufacturing.com or you may call us at 617.909.6275.
Create a Job Alert
Interested in building your career at Re:Build Manufacturing? Get future opportunities sent straight to your email.
Apply for this job
*
indicates a required field
