Back to jobs

Senior AI/ML Test and Evaluation Engineer

Washington, DC Metro | Denver, CO Metro | Colorado Springs, CO

Who We Are

Every organization runs on intelligence: years of accumulated knowledge, decisions, and context. As AI takes on more of that work, companies face a choice: rent that intelligence from vendors who keep the data, the context, and the results, or own it.

OpenTeams exists to make ownership possible.

Founded by Travis Oliphant, creator of NumPy and SciPy, and built by people with deep roots across the open-source ecosystem, including NumPy, SciPy, PyTorch, and Jupyter, we help enterprises and governments build AI they control, govern, and evolve themselves.

If that sounds like your kind of work, we'd like to meet you.

Senior AI/ML Test and Evaluation Engineer

Location: Washington, DC; Denver, CO; or Colorado Springs, CO preferred (hybrid). Highly qualified candidates outside these locations may also be considered.  

Work Authorization: U.S. citizenship required

Clearance: An active U.S. security clearance is strongly preferred. Candidates without an active clearance may be considered for unclassified work but must be eligible to obtain and maintain a clearance. 

Salary Range: $145,000–$250,000 USD, dependent on experience level and location

About the Role

We're looking for a Senior AI/ML Test and Evaluation Engineer to build and operate the benchmarking and evaluation capability at the core of an AI platform. This is a role for someone who is more interested in what a model gets wrong than in what it gets right.

You build the evaluation harnesses — automated metrics paired with structured human expert judgment, applied to candidate models and to the agentic workflows built on top of them. You develop repeatable methodologies for comparing performance against current operational baselines, which means the comparison holds up when someone runs it again in six months with a different model. And you document the limitations and surface the failure modes that matter, including the ones nobody asked about.

Your reports go to senior stakeholders and inform decisions about which capabilities are ready to field. That's the weight of the job: a benchmark that looks good and hides a failure mode is worse than no benchmark at all, and you're the check against that.

This is hands-on engineering on an open-source toolchain.

This position is contingent upon contract award. Travel of up to 15% may be required, primarily to Government facilities and between company locations.

Key Responsibilities

  • Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows
  • Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment
  • Curate and recommend candidate benchmarks based on mission needs and document the provenance of ground-truth and reference data
  • Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes
  • Define and contribute to common standards for benchmark expression, ingestion, and reporting
  • Support partner organizations and vendors as they integrate their capabilities with shared evaluation standards
  • Build lightweight expert-scoring workflows and measure inter-reviewer agreement for judgment-based evaluations
  • Participate in structured feedback sessions with mission end users and incorporate findings into the platform and evaluation methodology
  • Develop reference notebooks and example workflows that enable data-science-capable analysts to run, interpret, and extend evaluations
  • Document technical approaches, evaluation results, and key decisions for Government stakeholders and internal teams

Required Skills & Experience

  • U.S. citizenship and eligibility to obtain and maintain a U.S. security clearance
  • 6+ years of software engineering or machine learning engineering experience, including 3+ years evaluating, benchmarking, or deploying ML models in production or applied research environments
  • Strong Python proficiency in a machine learning or data science context
  • Hands-on experience with common ML frameworks and tooling, such as PyTorch and the Hugging Face ecosystem
  • Experience developing or using model evaluation harnesses, benchmark suites, or test and evaluation frameworks
  • Experience designing evaluation metrics and applying appropriate statistical rigor when interpreting and reporting results
  • Experience building repeatable and auditable evaluation pipelines with documented data provenance
  • Experience evaluating large language models or agentic workflows using task-based, metric-based, or judgment-based scoring
  • Strong written communication skills, including the ability to clearly explain evaluation methodologies, results, limitations, and failure modes to technical and nontechnical stakeholders
  • Ability to work effectively in an evolving environment and translate mission needs into practical evaluation approaches
  • Bachelor’s degree in computer science, mathematics, engineering, or a related field, or equivalent practical experience

Nice to Have

  • Active U.S. security clearance
  • Prior AI/ML evaluation or test and evaluation experience supporting the Department of Defense, Intelligence Community, or another federal customer
  • Experience designing human-in-the-loop evaluations, measuring inter-rater reliability, or facilitating structured expert adjudication
  • Experience defining or implementing benchmark interchange formats or evaluation standards used across multiple organizations
  • Familiarity with intelligence analysis workflows or other high-stakes analytical domains
  • Experience working directly with Government stakeholders, mission users, or external technical partners
  • Contributions to open-source machine learning, benchmarking, or evaluation projects

What We Offer

  • Medical, Dental & Vision – 100% paid for employees, 75% for dependents
  • 401(k) Match – Up to 5% with full vesting after 2 years
  • Unlimited PTO – With a required minimum of 15 days off annually
  • Fully Remote Setup – Includes up to $3,000 equipment reimbursement
  • Continuous Education –  Includes up to $500 reimbursement
  • Disability & Life Insurance – 100% employer-paid
  • HSA & FSA Options – With monthly HSA contributions from OpenTeams

 

Grow With Us

At OpenTeams, growth isn’t just about the company—it’s about you.

We believe the best careers are built at the edge of your potential. That is where new tools, ideas, and technologies change the world. Here, you’ll work alongside pioneers of AI, solving problems that matter: making AI more transparent, more ethical, and more empowering. As your skills grow, our career framework provides a pathway and recognition of that increased impact.

Opportunities aren’t limited by geography. You’ll collaborate with global experts, contribute to open source projects that power the world’s technology, and stretch your skills daily.  That global perspective and diversity makes our solution more universal and robust.  We are committed to continuing to celebrate diversity on our team.

Supported people are successful people.  We offer 100% employer paid medical premiums for employees and self-managed PTO with a minimum time off requirement, so that our teams are able to do their best work.

We invest  in curiosity, creativity, and ownership. That means you’ll be trusted to boldly innovate, supported to learn fast, and celebrated for successful collaboration.

Commitment to diversity, equity, inclusion, and belonging

OpenTeams understands that valuing diverse creative practices and forms of knowledge is crucial to and enriches the company’s core mission. We encourage applications from everyone, including members of all equity-seeking communities, such as (but certainly not limited to) women, racialized and Indigenous persons, disabled people, persons of all sexual orientations, gender identities and expressions.

We are an equal opportunity employer - all qualified applicants will receive equal consideration for recruitment, interviews, employment, training, compensation, promotion, and related activities. We do not discriminate based on race, religion, gender, gender identity, gender expression, color, national origin, pregnancy, ancestry, domestic partner status, disability, sexual orientation, age, genetic predisposition, medical condition, marital status, citizenship status, military or veteran status, or any other basis covered by applicable laws. OpenTeams will not tolerate discrimination or harassment based on these characteristics or any other unlawful behavior, conduct, or purpose.

 

Apply for this job

*

indicates a required field

Phone
Resume/CV

Accepted file types: pdf, doc, docx, txt, rtf

Cover Letter

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...
Which of the following areas do you have hands-on professional experience with? *

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in OpenTeams’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...

Voluntary Self-Identification of Disability

Form CC-305
Page 1 of 1
OMB Control Number 1250-0005
Expires 07/31/2029

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

  • Alcohol or other substance use disorder (not currently using drugs illegally)
  • Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
  • Blind or low vision
  • Cancer (past or present)
  • Cardiovascular or heart disease
  • Celiac disease
  • Cerebral palsy
  • Deaf or serious difficulty hearing
  • Diabetes
  • Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
  • Epilepsy or other seizure disorder
  • Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
  • Intellectual or developmental disability
  • Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
  • Missing limbs or partially missing limbs
  • Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
  • Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
  • Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
  • Partial or complete paralysis (any cause)
  • Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
  • Short stature (dwarfism)
  • Traumatic brain injury
Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.


We use Greenhouse’s AI-powered Talent Matching tool to compare your application against our job requirements.

Learn more