NextRaise Logo
JobsTrackerResumes
Job MatchATS
ResumesJobsProfile
Jobs / AI / ML Researcher in United States of America
3 days ago
Apply with autofill
Apply with autofill
Vettoai·3 days ago
3 days ago

AI Research

LDN, United States of AmericaRemoteMid · 2-5 years

Sign up free to see how well your resume matches this role.

About this role

AI Researcher

About Vetto

Vetto builds the infrastructure for next-generation AI training data. We partner with the world’s top AI labs to push the frontier of what models can do — by designing, collecting, and delivering the highest-quality training data for the hardest problems in AI.

AI Researchers at Vetto figure out how to teach models new capabilities and how to tell whether it worked. You’ll design datasets and evaluations, run experiments, study model failures, and turn what you learn into better data and training strategies.

Most of our work is evaluation-driven, with some fine-tuning and post-training used to validate research hypotheses. You won’t be expected to build the infrastructure alone: you’ll work with an engineering team that builds the platforms and tools behind the research. You should still be comfortable writing code, analyzing data, and moving your own experiments forward.

Why Vetto?

  • End-to-end ownership. You’ll take research questions from an initial idea to a dataset, benchmark, experiment, and useful conclusion.
  • Direct collaboration with top AI labs. Your work will help frontier teams understand their models and decide what data or training approach to try next.
  • Work that gets used. Your research can become training data, evaluations, technical reports, open benchmarks, and published papers.
  • A wide range of hard problems. Depending on the project, you might work on agents, model evaluations, human or synthetic data, failure analysis, or post-training.
  • Fast, flat team. You’ll ship experiments in days, not quarters.

What You’ll Do

  • Design and run experiments to understand model behavior, test new ideas, and determine whether a dataset or training intervention actually works.
  • Use fine-tuning and post-training experiments when useful to validate research hypotheses and measure improvements in the capabilities we care about.
  • Create datasets, benchmarks, agent environments, rubrics, and evaluation methods for problems that are not well measured today.
  • Work directly with AI labs and enterprise partners to turn open-ended model-development goals into concrete research and data projects.
  • Dig into model outputs and trajectories to understand systematic failure modes and separate model limitations from problems in the data, grader, or evaluation setup.
  • Write analysis code and build lightweight tools or prototypes for your own work, with support from engineers when a workflow needs to become reliable infrastructure.
  • Share what you learn through clear technical reports, internal discussions, benchmarks, and research papers.

What We’re Looking For

Must-haves:

  • Experience doing empirical research or similarly open-ended technical work with machine learning systems.
  • Hands-on experience designing and running agentic systems for complex work, such as coding agents and automated research pipelines.
  • Good experimental judgment: you can form a hypothesis, choose useful baselines and metrics, run a careful experiment, and make sense of noisy results.
  • Strong Python and data analysis skills, plus enough software engineering ability to prototype your own research workflows.
  • The independence to take an ambiguous question about a model and turn it into a dataset, benchmark, experiment, or other concrete contribution.
  • Rigorous analytical thinking — you look at the underlying evidence, notice subtle failure modes, and question results that seem too simple or too good to be true.
  • Excellent written and verbal communication with researchers, engineers, customers, and non-specialists.
  • Comfort with ambiguity, fast iteration, changing priorities, and occasional forward-deployed work with partner teams.

Nice to have:

  • Research or industry experience in LLM evaluation, post-training, agents, human or synthetic data, alignment, or AI safety.
  • Experience designing datasets or benchmarks, validating annotation quality, or working with large collections of model outputs and trajectories.
  • Published research, technical writing, open-source contributions, or strong independent projects.
  • Experience building or managing human feedback pipelines.
  • A background in cognitive science, linguistics, philosophy, statistics, or another field that helps you think clearly about evaluation.
  • Experience at an AI lab, AI data company, or in a partner-facing research role.

We care more about evidence of excellent work than any single credential. A graduate degree, publications, lab experience, open-source work, independent research, and production engineering experience are all useful signals, but none is a requirement on its own.

What We Offer

  • Competitive compensation and a stock option plan.
  • Global, remote-first work with flexibility and periodic in-person on-sites every few months.
  • Direct collaboration with cutting-edge AI labs on evaluation, data, and post-training challenges.
  • High ownership and fast career growth in a founding-era team.

Location: Global / Remote, with periodic in-person on-sites every few months

Reports to: Co-founders

H1B sponsor likely
AI tools
Apply faster with autofillThe NextRaise extension autofills your application in one click.Get the extension

Similar jobs

  • Research Scientist - VLM Pretraining at epsilon-healthSan Francisco, United States of America
  • Research Scientist - Vision Foundation Models at epsilon-healthSan Francisco, United States of America
  • Research Scientist - Post-training / RL at epsilon-healthSan Francisco, United States of America
  • Research Engineer - Data Quality & Evals at epsilon-healthSan Francisco, United States of America
  • Pre-training Research Engineer at sciforiumSan Francisco, United States of America
  • Research Engineer - Model Evaluation & MLOps at sciforiumSan Francisco, United States of America