Back to jobs
New

AI Evaluation Lead

United States

 

The Mission

We live in a paradox: AI is accelerating the world’s capabilities, yet the average person feels more financially precarious than ever. Inflation is rising, wages are stagnant, and the traditional “retirement” model is broken. We aren’t building another chatbot. We are building the Financial Answer Machine, an intelligent guide designed to help people navigate a new financial reality.
 
Underpinned by a proprietary financial system, we are turning “average” advice into personalized, multi-modal financial power. We have closed an over-subscribed seed round and are looking for founding team members to help us build a bridge between the intelligence of AI and the rigid accuracy required for financial freedom. This is a rare opportunity to join at Day Zero and architect a business designed for outsized impact and massive scale.

The Role

Our client is hiring an AI Evaluation Lead to own how we measure the quality of AI-generated financial advice. Getting the advice right matters. A bad output here has real consequences for real people, and this role owns making sure we catch it.
 
You will work with an AI-generated test case library and automated scoring infrastructure that is already in place. Your job is to make sure we are measuring the right things, interpreting what the results are telling us, and determining what needs to change to keep the system performing well as it scales. Evaluation complexity grows with the platform, and this role grows with it.
 
This is not a monitoring and reporting role. It requires genuine judgment about AI system behavior, advice quality, and what the data is and is not capturing. You will report to our Head of Revenue & Compliance, work closely with the AI/ML team and founders, and partner with subject matter experts who provide domain judgment on complex or ambiguous cases. But you need enough personal finance literacy to make first-pass quality assessments independently and know when to escalate.

What You’ll Do (The Day-to-Day)

  • Define and validate the evaluation set: what cases we should be testing, whether coverage is sufficient across domains, and where the current framework has gaps.
  • Analyze scoring results to identify highest-frequency case types, patterns in what is performing well versus poorly, and anomalies that warrant closer review.
  • Assess whether current measures are detecting the right failure modes or whether new measures are needed.
  • Review flagged cases and make judgment calls on what the results mean and what should be done about them, drawing on both data and domain knowledge.
  • Own the criteria and calibration for when human review is triggered: defining what rises to that level, what does not, and ensuring the threshold stays well-calibrated as the platform scales.
  • Partner with subject matter experts on cases that require deeper domain judgment, and incorporate their input into evaluation design.
  • Ensure evaluation coverage keeps pace with new domain additions and model changes before they ship.
  • Translate findings into specific, actionable recommendations for the AI/ML team on what needs to change in the system.
  • Evolve the evaluation framework as the system grows, new domains are added, and user patterns shift.
 

What We’re Looking For

You have worked on AI or ML system quality in a context where outputs had real stakes. You think analytically about what data is and is not telling you. You are comfortable making judgment calls in ambiguous situations rather than waiting for the answer to be obvious. You have enough AI/ML fluency to reason about why a system is producing what it is producing, not just whether the output looks right.
 
You bring enough personal finance literacy to read an advice response and have a genuine opinion about whether it is directionally sound. You do not need formal credentials or deep expertise across every domain the system covers—you will partner with subject matter experts for the complex judgment calls. What matters is that your review is substantive rather than mechanical, and that you can have an informed conversation with those experts about what you are seeing in the data.
 
  • Fluency with how LLM-based systems behave in production, including output variance, failure modes, and the limits of automated scoring.
  • Ability to assess whether an eval framework is measuring the right things, not just whether it is running correctly.
  • Comfortable working with behavioral and interaction data to surface patterns and quality signals.
  • Familiarity with evaluation and observability tooling.

Backgrounds that tend to fit:

  • Model evaluation or QA on a consumer-facing AI product, particularly in a regulated or high-stakes context.
  • Model risk or validation with LLM or generative AI exposure.
  • Data science or analytics with ownership of production AI system quality.
  • Operations quality control built around AI- or ML-generated outputs.
  • Financial services or fintech product roles where you developed both analytical depth and personal finance domain familiarity.

This is probably not the right role for you if: 

  • Your background is primarily in building models rather than evaluating what they produce
  • Personal finance is entirely unfamiliar territory. You do not need to be an expert, but you need enough baseline literacy to assess whether advice is reasonable and to work productively with the SMEs who provide deeper domain judgment
  • You are looking for a well-defined role with stable processes. The framework is in place but evolving it is a core part of the job
  • You default to manual review rather than thinking systematically about what should be automated and what requires human judgment

How we work

We are a fully remote, distributed team. Periodic in-person get-togethers will be integral to our operating cadence. We’re adults who prioritize outcomes and output over set schedules. We value clear writing, high ownership, fast iteration, direct communication, and thoughtful async collaboration.
 
As an early team member, you should expect broad ownership, frequent context shifts, and a high degree of autonomy. You will help shape not just the product, but also the technical standards and operating cadence of the company.

Compensation

Salary: 120-140k, plus early-stage option equity.
Final compensation will depend on level, experience, location, and scope of responsibility.
This role is open to candidates based in the United States.
 

AI Interview

We expect a high volume of applications for this role. To help candidates showcase more than what's on their resume, you'll have the opportunity to complete an AI interview as part of the application process.

As an AI-first company, we embrace AI throughout the hiring process and are excited to meet candidates who are equally curious about and enthusiastic about the technology. This interview is your chance to demonstrate your experience, communication skills, and potential beyond your resume.

 

 

Apply for this job

*

indicates a required field

Phone
Resume/CV*

Accepted file types: pdf, doc, docx, txt, rtf


Select...
Select...
Select...
Select...

Voluntary Self-Identification

For government reporting purposes, we ask candidates to respond to the below self-identification survey. Completion of the form is entirely voluntary. Whatever your decision, it will not be considered in the hiring process or thereafter. Any information that you do provide will be recorded and maintained in a confidential file.

As set forth in elly’s Equal Employment Opportunity policy, we do not discriminate on the basis of any protected group status under any applicable law.

Select...
Select...
Race & Ethnicity Definitions

If you believe you belong to any of the categories of protected veterans listed below, please indicate by making the appropriate selection. As a government contractor subject to the Vietnam Era Veterans Readjustment Assistance Act (VEVRAA), we request this information in order to measure the effectiveness of the outreach and positive recruitment efforts we undertake pursuant to VEVRAA. Classification of protected categories is as follows:

A "disabled veteran" is one of the following: a veteran of the U.S. military, ground, naval or air service who is entitled to compensation (or who but for the receipt of military retired pay would be entitled to compensation) under laws administered by the Secretary of Veterans Affairs; or a person who was discharged or released from active duty because of a service-connected disability.

A "recently separated veteran" means any veteran during the three-year period beginning on the date of such veteran's discharge or release from active duty in the U.S. military, ground, naval, or air service.

An "active duty wartime or campaign badge veteran" means a veteran who served on active duty in the U.S. military, ground, naval or air service during a war, or in a campaign or expedition for which a campaign badge has been authorized under the laws administered by the Department of Defense.

An "Armed forces service medal veteran" means a veteran who, while serving on active duty in the U.S. military, ground, naval or air service, participated in a United States military operation for which an Armed Forces service medal was awarded pursuant to Executive Order 12985.

Select...

Voluntary Self-Identification of Disability

Form CC-305
Page 1 of 1
OMB Control Number 1250-0005
Expires 04/30/2026

Why are you being asked to complete this form?

We are a federal contractor or subcontractor. The law requires us to provide equal employment opportunity to qualified people with disabilities. We have a goal of having at least 7% of our workers as people with disabilities. The law says we must measure our progress towards this goal. To do this, we must ask applicants and employees if they have a disability or have ever had one. People can become disabled, so we need to ask this question at least every five years.

Completing this form is voluntary, and we hope that you will choose to do so. Your answer is confidential. No one who makes hiring decisions will see it. Your decision to complete the form and your answer will not harm you in any way. If you want to learn more about the law or this form, visit the U.S. Department of Labor’s Office of Federal Contract Compliance Programs (OFCCP) website at www.dol.gov/ofccp.

How do you know if you have a disability?

A disability is a condition that substantially limits one or more of your “major life activities.” If you have or have ever had such a condition, you are a person with a disability. Disabilities include, but are not limited to:

  • Alcohol or other substance use disorder (not currently using drugs illegally)
  • Autoimmune disorder, for example, lupus, fibromyalgia, rheumatoid arthritis, HIV/AIDS
  • Blind or low vision
  • Cancer (past or present)
  • Cardiovascular or heart disease
  • Celiac disease
  • Cerebral palsy
  • Deaf or serious difficulty hearing
  • Diabetes
  • Disfigurement, for example, disfigurement caused by burns, wounds, accidents, or congenital disorders
  • Epilepsy or other seizure disorder
  • Gastrointestinal disorders, for example, Crohn's Disease, irritable bowel syndrome
  • Intellectual or developmental disability
  • Mental health conditions, for example, depression, bipolar disorder, anxiety disorder, schizophrenia, PTSD
  • Missing limbs or partially missing limbs
  • Mobility impairment, benefiting from the use of a wheelchair, scooter, walker, leg brace(s) and/or other supports
  • Nervous system condition, for example, migraine headaches, Parkinson’s disease, multiple sclerosis (MS)
  • Neurodivergence, for example, attention-deficit/hyperactivity disorder (ADHD), autism spectrum disorder, dyslexia, dyspraxia, other learning disabilities
  • Partial or complete paralysis (any cause)
  • Pulmonary or respiratory conditions, for example, tuberculosis, asthma, emphysema
  • Short stature (dwarfism)
  • Traumatic brain injury
Select...

PUBLIC BURDEN STATEMENT: According to the Paperwork Reduction Act of 1995 no persons are required to respond to a collection of information unless such collection displays a valid OMB control number. This survey should take about 5 minutes to complete.