
Senior DevOps/Platform Engineer
About Goods & Services
Goods & Services is a product design and engineering company.
We solve mission-critical challenges for some of the world’s largest enterprises, with deep expertise in highly regulated industries—including life sciences and financial services. Our design-led approach allows us to apply cutting-edge capabilities in AI, Data and Hardware Engineering to companies of any size.
Headquartered in the United States, we operate regional development centers in Mexico and the United Kingdom. This global footprint—anchored by our nearshore model—enables us to deliver at scale with the speed, efficiency, and cultural alignment our clients expect.
At Goods and Services, we value diversity and are committed to creating an inclusive workplace where everyone can thrive. We are proud to be an Equal Opportunity Employer and welcome qualified applicants from all backgrounds.
About the job
Goods & Services is looking for a Senior DevOps/Platform Engineer to build and maintain the infrastructure supporting reliable application delivery, deployment, security, and production operations. This role will focus on CI/CD automation, feature management, observability, infrastructure reliability, and secure platform practices while enabling teams to deliver and validate changes confidently at scale.
What you’ll do:
- Provision and maintain development, test, and production environments, including access management, CI/CD pipelines, source control, and infrastructure needed to support application delivery.
- Implement and maintain platform security and observability capabilities, including WAF, rate limiting, distributed tracing, monitoring, and related infrastructure.
- Partner with engineering and technical leadership to define and validate traffic-ramp strategies, release criteria, and minimum stabilization periods for production deployments.
- Build and maintain feature-flag infrastructure that enables controlled releases, incremental migrations, and safe rollback of application components.
- Monitor and maintain event-driven infrastructure, including alerts for failed events, dead-letter queues, schema lifecycle issues, and data backup or archival processes.
- Develop and maintain service reliability and integration monitoring, including dashboards, alerting, and automated notifications for failures across internal and third-party services.
- Create and maintain operational runbooks and incident-response procedures for common platform, integration, workflow, and deployment issues.
- Support staged production rollouts and traffic validation, ensuring new services are stable before fully transitioning traffic and safely retiring legacy components.
What you’ll need:
- Solid AWS CI/CD experience, using tools such as GitHub Actions, AWS CodePipeline, or equivalent, in multi-service cloud environments.
- Hands-on experience with AWS observability and monitoring, including CloudWatch, AWS X-Ray, logging, alerting, dashboards, and paging workflows.
- Experience securing and hardening AWS WAF and API Gateway, including managed security rules, rate limiting, access controls, and traffic protection.
- Experience implementing and managing feature-flag platforms to support controlled releases, incremental migrations, parallel-run scenarios, and safe rollbacks.
- Experience with infrastructure and environment management, including provisioning, configuration, access management, and supporting development, test, and production environments.
- Experience monitoring event-driven and asynchronous systems, including EventBridge, message queues, dead-letter queues, event failures, and alerting mechanisms.
- Experience implementing service reliability and integration monitoring, including circuit-breaker monitoring, failure detection, dashboards, and automated alerting for third-party services.
- Solid incident response and operational readiness experience, including creating runbooks, troubleshooting production issues, and defining recovery procedures.
- Experience supporting production deployments and controlled traffic rollouts, including defining ramp-up criteria, soak periods, validation metrics, and rollback strategies.
- Comfortable collaborating with engineering and technical leadership to establish deployment standards, operational processes, and production-readiness criteria.
Apply for this job
*
indicates a required field