
AI Research
AI Researcher
About Vetto
Vetto builds the infrastructure for next-generation AI training data. We partner with the world’s top AI labs to push the frontier of what models can do — by designing, collecting, and delivering the highest-quality training data for the hardest problems in AI.
AI Researchers at Vetto figure out how to teach models new capabilities and how to tell whether it worked. You’ll design datasets and evaluations, run experiments, study model failures, and turn what you learn into better data and training strategies.
Most of our work is evaluation-driven, with some fine-tuning and post-training used to validate research hypotheses. You won’t be expected to build the infrastructure alone: you’ll work with an engineering team that builds the platforms and tools behind the research. You should still be comfortable writing code, analyzing data, and moving your own experiments forward.
Why Vetto?
- End-to-end ownership. You’ll take research questions from an initial idea to a dataset, benchmark, experiment, and useful conclusion.
- Direct collaboration with top AI labs. Your work will help frontier teams understand their models and decide what data or training approach to try next.
- Work that gets used. Your research can become training data, evaluations, technical reports, open benchmarks, and published papers.
- A wide range of hard problems. Depending on the project, you might work on agents, model evaluations, human or synthetic data, failure analysis, or post-training.
- Fast, flat team. You’ll ship experiments in days, not quarters.
What You’ll Do
- Design and run experiments to understand model behavior, test new ideas, and determine whether a dataset or training intervention actually works.
- Use fine-tuning and post-training experiments when useful to validate research hypotheses and measure improvements in the capabilities we care about.
- Create datasets, benchmarks, agent environments, rubrics, and evaluation methods for problems that are not well measured today.
- Work directly with AI labs and enterprise partners to turn open-ended model-development goals into concrete research and data projects.
- Dig into model outputs and trajectories to understand systematic failure modes and separate model limitations from problems in the data, grader, or evaluation setup.
- Write analysis code and build lightweight tools or prototypes for your own work, with support from engineers when a workflow needs to become reliable infrastructure.
- Share what you learn through clear technical reports, internal discussions, benchmarks, and research papers.
What We’re Looking For
Must-haves:
- Experience doing empirical research or similarly open-ended technical work with machine learning systems.
- Hands-on experience designing and running agentic systems for complex work, such as coding agents and automated research pipelines.
- Good experimental judgment: you can form a hypothesis, choose useful baselines and metrics, run a careful experiment, and make sense of noisy results.
- Strong Python and data analysis skills, plus enough software engineering ability to prototype your own research workflows.
- The independence to take an ambiguous question about a model and turn it into a dataset, benchmark, experiment, or other concrete contribution.
- Rigorous analytical thinking — you look at the underlying evidence, notice subtle failure modes, and question results that seem too simple or too good to be true.
- Excellent written and verbal communication with researchers, engineers, customers, and non-specialists.
- Comfort with ambiguity, fast iteration, changing priorities, and occasional forward-deployed work with partner teams.
Nice to have:
- Research or industry experience in LLM evaluation, post-training, agents, human or synthetic data, alignment, or AI safety.
- Experience designing datasets or benchmarks, validating annotation quality, or working with large collections of model outputs and trajectories.
- Published research, technical writing, open-source contributions, or strong independent projects.
- Experience building or managing human feedback pipelines.
- A background in cognitive science, linguistics, philosophy, statistics, or another field that helps you think clearly about evaluation.
- Experience at an AI lab, AI data company, or in a partner-facing research role.
We care more about evidence of excellent work than any single credential. A graduate degree, publications, lab experience, open-source work, independent research, and production engineering experience are all useful signals, but none is a requirement on its own.
What We Offer
- Competitive compensation and a stock option plan.
- Global, remote-first work with flexibility and periodic in-person on-sites every few months.
- Direct collaboration with cutting-edge AI labs on evaluation, data, and post-training challenges.
- High ownership and fast career growth in a founding-era team.
Location: Global / Remote, with periodic in-person on-sites every few months
Reports to: Co-founders
Apply for this job
*
indicates a required field