AI Reliability Engineer (HPC)
Overview
As Microsoft continues to push the boundaries of AI, we are on the lookout for passionate individuals to work with us on the most interesting and challenging AI questions of our time. Our vision is bold and broad — to build systems that have true artificial intelligence across agents, applications, services, and infrastructure. It’s also inclusive: we aim to make AI accessible to all — consumers, businesses, developers — so that everyone can realize its benefits.
We’re looking for an experienced Member of Technical Staff – AI Reliability Engineer to join our High Performance Computing (HPC) infrastructure team. In this role, you’ll blend software engineering and systems engineering to keep our large-scale distributed AI infrastructure reliable and efficient. You’ll ensure that AI systems stay efficient and reliable with very high uptimes.
Responsibilities
- Reliability & Availability: Ensure uptime, resiliency, and fault tolerance of HPC clusters powering MAI model training and inference.
- Observability: Design and maintain monitoring, alerting, and logging systems to provide real-time visibility into all aspects of HPC systems including GPU, clusters, storage and networking.
- Automation & Tooling: Build automation for deployments, incident response, scaling, and failover in CPU+GPU environments.
- Incident Management: Lead on-call rotations, troubleshoot production issues, conduct blameless postmortems, and drive continuous improvements.
- Security & Compliance: Ensure data privacy, compliance, and secure operations across model training and serving environments.
- Collaboration: Partner with ML engineers and platform teams to improve developer experience and accelerate research-to-production workflows.
- Embody our Culture and Values.
Qualifications
- Bachelor’s Degree in Computer Science, or related technical discipline AND 4+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering
- OR equivalent experience
- Master’s Degree in Computer Science, or related technical discipline AND 2+ years technical experience in Site Reliability Engineering, DevOps, or Infrastructure Engineering
- OR equivalent experience
-
Experience with Kubernetes, Docker, container orchestration, and CI/CD pipelines for ML training or inference workloads.
-
Experience with public cloud platforms such as Azure, AWS, or GCP, including infrastructure-as-code.
-
Experience with monitoring and observability tools such as Grafana, Datadog, or OpenTelemetry.
-
Programming or scripting experience in Python, Go, or Bash.
-
Experience with distributed systems, networking, storage, and high-performance computing (HPC).
-
Experience operating GPU clusters and workload schedulers for ML/AI workloads.
-
Experience with ML training or inference pipelines.
-
Experience with capacity planning and cost optimization for GPU-based infrastructure.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Apply for this job
*
indicates a required field
