Back to jobs
New

Software Engineer, Model API Infra

Mountain View, Redmond

About the Role

As a Member of Technical Staff, Model API Infra, you will own the production API and agentic platform for MAI models. Combining strong software engineering with a deep understanding of how models are trained, evaluated, and served, you will build harness infrastructure ensuring training and production parity, automate the model deployment lifecycle, and create cross-stack observability and tooling to run these services reliably at scale. You will collaborate closely with post-training, evaluation, Foundry, Copilot, and product teams, to continuously deliver new model capabilities to production surfaces.

Job Overview

Model API and Agentic Platform 

  • Design, implement, and operate scalable API services for text, multimodal, and emerging MAI models, with strong compatibility, versioning, reliability, and performance guarantees. 
  • Build API and systems that maintain consistent model, tool, and harness behavior across training, evaluation, and production serving. 
  • Partner with Foundry, Copilot platforms, inference, research, eval, and product teams to provide consistent interfaces and enable rapid model iteration and delivery. 
  • Build business, API, and model system telemetry that enables cross-system debugging, automated investigations, and data-driven insights. 

Model Serving Automation 

  • Build a closed-loop system that automates model deployment and validation. 
  • Integrate functional, safety, quality, and performance evaluations into deployment decisions; debug eval regressions across the stack, from checkpoint and serving config to API and product. 
  • Design and build autoscaling for model pools using latency SLOs, utilization, model characteristics, and capacity constraints. 

Required Qualifications

  • Bachelor's Degree in Computer Science or related technical field AND 5+ years of technical engineering experience with coding in languages including Python, C++, or similar OR equivalent experience.
  • Experience with LLM inference/serving systems (vLLM, SGLang, TensorRT-LLM, or custom inference engines).
  • Strong systems/debugging skills and comfort analyzing performance on GPU clusters.
  • Experience with distributed serving/inference concepts (TP/PP/DP, KV cache, speculative decoding).
  • Experience with deployment automation, orchestration such as Kubernetes, and cloud platforms. 

Preferred Qualifications

  • Master's Degree AND 5+ years of relevant experience OR equivalent experience. 
  • Experience deploying models to production serving platforms (Azure ML/Foundry, Kubernetes, containerized deployments). 
  • Familiarity with multi-adapter serving, disaggregated prefill/decode, or large-scale KV cache management. 
  • Experience with observability tools (Grafana, Prometheus or equivalent, Kusto) and analyzing production telemetry. 
  • Experience building evaluation/benchmarking harnesses for LLMs. 
  • Knowledge of GPU topology, memory management, and modern accelerator hardware. 


Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

 

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...