Back to jobs
New

Software Engineer, Model Serving System

Mountain View, Redmond

About the Role

We're
hiring a Software Engineer to join a fast
‑moving, high‑ownership team building the next‑generation large language model (LLM) serving system. This role builds and optimizes the end-to-end path to get Microsoft AI (MAI)'s models running reliably in production through Microsoft Azure cloud platform. We work at the intersection of distributed model serving systems, inference engine optimization and integration, benchmarking, and evaluation to make MAI models shippable, measurable, and evaluable.

Job Overview

Model Serving & Integration 

  • Partner with engine teams to integrate, optimize, and deploy models to production serving environments with high reliability. 
  • Help implement and evaluate production serving features (e.g., content provenance/watermarking) while minimizing latency impact. 
  • Profile and optimize for production SLOs (TTFT, TPOT, throughput, QPS) on modern GPU hardware and distributed serving system topology. 
  • Work across platform, API, and research teams to streamline the path from model release to production availability. 

Model Fine-tuning & Adaptation Support 

  • Support serving of adapted models (including multi-LoRA scenarios) with careful attention to memory and caching behavior. 
  • Help build tooling to package and validate fine-tuned models for production deployment. 
  • Improve the reliability and repeatability of deployment workflows. 

Benchmarking, Evaluation & Observability 

  • Design and run benchmarks using real usage patterns to measure serving performance under realistic conditions. 
  • Make production endpoints evaluable, identify bottlenecks in evaluation workflows, and improve their speed and reliability. 
  • Build and improve dashboards and telemetry to track utilization, latency, cache efficiency, and SLA compliance. 
  • Translate measurement insights into concrete optimizations with quantifiable impact. 

Required Qualifications

  • Bachelor's Degree in Computer Science or related technical field AND 5+ years of technical engineering experience with coding in languages including Python, C++, or similar OR equivalent experience. 
  • Experience with LLM inference/serving systems (vLLM, SGLang, TensorRT-LLM, or custom inference engines). 
  • Strong systems/debugging skills and comfort analyzing performance on GPU clusters. 
  • Experience with distributed serving/inference concepts (TP/PP/DP, KV cache, speculative decoding). 

Preferred Qualifications

  • Master's Degree AND 5+ years of relevant experience OR equivalent experience. 
  • Experience deploying models to production serving platforms (Azure ML/Foundry, Kubernetes, containerized deployments). 
  • Familiarity with multi-adapter serving, disaggregated prefill/decode, or large-scale KV cache management. 
  • Experience with observability tools (Grafana, Prometheus or equivalent, Kusto) and analyzing production telemetry. 
  • Experience building evaluation/benchmarking harnesses for LLMs. 
  • Knowledge of GPU topology, memory management, and modern accelerator hardware. 


    Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.

    Software Engineering IC5 - The typical base pay range for this role across the U.S. is USD $142,800 - $274,800 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $188,000 - $304,200 per year

 

 

Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay


This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.

Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.

 

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...