Production Safety
About Microsoft AI
Microsoft AI is building AI systems and products that empower people’s lives. Our work is driven by a community of brilliant, interdisciplinary minds working across frontier model development, product engineering, and responsible AI. Within Microsoft AI, the Safety team develops the training methods, evaluations, runtime safeguards, monitoring, and infrastructure needed to make advanced AI systems safer, more reliable, and more useful. Our work spans text, multimodal, and agentic systems and is developed in close partnership with other Research teams, Production Inference, Security, Responsible AI, Microsoft product organizations, and external partners and customers.
About the Role
We are looking for a Member of Technical Staff to build, deploy, and operate the production systems that keep advanced AI products safe and reliable at global scale. You will work on safety-critical components in the inference path, including orchestration, model- and rules-based safeguards, configuration, telemetry, fail-safe behavior, and rollout mechanisms. You will also build production monitoring for large-scale training and evaluation runs, ensuring that pipelines, compute, and safety signals remain healthy and reliable.
This role is a strong fit for a software engineer with experience in high-scale cloud and distributed systems and balancing product safety and quality with latency, availability, and cost.
Responsibilities
- Design, build, deploy, and operate safety-critical services in the production inference path for text, multimodal, and agentic AI systems.
- Integrate model-based classifiers, policy engines, and other guardrails with model APIs and serving platforms.
- Build containerized services on Kubernetes, with automated CI/CD, staged regional rollout, rollback, and failover.
- Define and meet service-level objectives for availability, latency, throughput, correctness, and cost, including capacity planning and autoscaling.
- Build monitoring and debugging capabilities that make safety decisions and service health observable, and detect failures, stalls, and regressions in large-scale training and evaluation runs.
- Validate launches across development, staging, and production, and contribute to on-call, incident response, and root-cause analysis.
- Partner with Production Inference, Security, Privacy, Responsible AI, and product teams to resolve dependencies and meet launch requirements.
- Contribute reusable libraries, testing standards, deployment runbooks, and operational practices for safety-critical infrastructure.
Required Qualifications
- Bachelor’s degree in Computer Science, Engineering, a related technical field, or equivalent practical experience.
- Strong software engineering skills in one or more production languages such as C++, C#, Java, Go, Rust, or Python.
- Experience designing, building, deploying, and operating distributed services or other large-scale production systems.
- Experience with containerized workloads, Kubernetes, and automated CI/CD pipelines.
- Knowledge of service reliability fundamentals, including observability, capacity planning, failure isolation, graceful degradation, and incident response.
- Ability to reason carefully about correctness, security, privacy, and failure modes in systems that affect end users.
- Ability to collaborate across engineering, machine learning, product, and policy teams and communicate system tradeoffs clearly.
Preferred Qualifications
- Experience with online inference platforms, model serving, API gateways, policy enforcement systems, or other latency-sensitive infrastructure.
- Experience deploying or operating machine learning models, training pipelines, or evaluation systems in production, including versioning, monitoring, failure detection, and rollback or recovery.
- Experience with Azure technologies.
- Experience with observability tooling and capacity planning.
- Familiarity with or interest in AI safety, trust and safety, abuse prevention, content moderation, security engineering, privacy, or responsible AI.
Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Apply for this job
*
indicates a required field
