Safety Post-training
About Microsoft AI
Microsoft AI is building AI systems and products that empower people’s lives. Our work is driven by a community of brilliant, interdisciplinary minds working across frontier model development, product engineering, and responsible AI. Within Microsoft AI, the Safety team develops the training methods, evaluations, runtime safeguards, monitoring, and infrastructure needed to make advanced AI systems safer, more reliable, and more useful. Our work spans text, multimodal, and agentic systems and is developed in close partnership with other Research teams, Production Inference, Security, Responsible AI, Microsoft product organizations, and external partners and customers.
About the Role
We are looking for a Member of Technical Staff to help build the shared systems and methods for safety post-training of frontier text, coding, and agentic models. You will connect data generation and curation, annotation, model and agent evaluation, training experiments, and production feedback into reproducible workflows that improve model behavior at scale. This role is a strong fit for an AI researcher or engineer who combines rigorous evaluation with strong engineering judgment and enjoys creating reproducible experiments and production-ready code.
Responsibilities
- Design and build shared pipelines for safety data collection and generation, evaluation, experimentation, and model feedback loops.
- Develop rigorous evaluation suites, datasets, metrics, and graders for text, coding, and agentic models.
- Define durable schemas and APIs for policy taxonomies, annotations, grader outputs, datasets, tool trajectories, and evaluation results.
- Connect production signals, incidents, red-team findings, and partner feedback to evaluations, training data, model interventions, and deployment decisions.
- Partner with partner and product teams to improve model and system safety.
Required Qualifications
- Bachelor’s degree in Computer Science, Machine Learning, Statistics, Mathematics, Engineering, a related technical field, or equivalent practical experience.
- Strong programming skills in Python.
- Experience training, fine-tuning, evaluating, or researching large language models.
- Experience designing and analyzing rigorous model evaluations, experiments, benchmarks, or measurement systems.
- Knowledge of experimental design and metric selection.
- Experience working with large datasets and building reliable data, evaluation, training, or workflow pipelines.
- Experience with data quality, annotation design, and automated graders.
- Ability to turn ambiguous safety or policy questions into measurable model behaviors and communicate results clearly across research, engineering, product, and policy teams.
Preferred Qualifications
Experience in one or more of the following areas is preferred:
- Safety post-training methods such as supervised fine-tuning, preference optimization, reinforcement learning from human or AI feedback, reward modeling, process supervision, rejection sampling, or synthetic data generation.
- Reinforcement learning infrastructure, rollout systems, agentic training environments, verifiers, trajectory-level graders, or large-scale distributed model training.
- Evaluation or mitigation of jailbreaks, prompt injection, unsafe tool use, reward hacking, deception, harmful cyber behavior, over-refusal, or other adversarial model behaviors.
- Cybersecurity expertise in areas such as application or cloud security, vulnerability research, penetration testing, secure coding, incident response, identity, or security engineering, with the ability to translate realistic threats into training data and evaluations.
- Experience building or evaluating coding and computer-use agents, instrumented sandboxes or cyber ranges, security-focused verifiers, authorization controls, or defenses against prompt injection and unsafe tool use.
- Statistical methods such as power analysis, Bayesian analysis, resampling, hierarchical models, sequential testing, multiple-comparison correction, causal inference, or measurement science.
- Human-evaluation design, automated-grader calibration, rare or high-severity event measurement, benchmark contamination analysis, subgroup analysis, distribution shift, or evaluator drift.
- Data-generation, annotation, evaluation, experiment-tracking, or model-lifecycle platforms using workflow orchestration, distributed compute, cloud infrastructure, lineage, governance, or production telemetry.
- Pretraining data and objectives, model architecture, interpretability, robustness, privacy, security, trust and safety, or other upstream model-development and responsible AI work.
Software Engineering IC4 - The typical base pay range for this role across the U.S. is USD $119,800 - $234,700 per year. There is a different range applicable to specific work locations, within the San Francisco Bay area and New York City metropolitan area, and the base pay range for this role in those locations is USD $160,200 - $261,000 per year.
Certain roles may be eligible for benefits and other compensation. Find additional benefits and pay information here:
https://careers.microsoft.com/us/en/us-corporate-pay
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Apply for this job
*
indicates a required field
