Research Scientist/Engineer, Frontier Reasoning, DeepMind
Sign up free to see how well your resume matches this role.
What you'll do
- Operate across the full research-and-engineering lifecycle of frontier reasoning and agentic systems.
- Work on unsolved problems in agentic reasoning, turning early exploratory prototypes into hardened production features for Gemini releases.
- Architect and optimize distributed post-training pipelines and agent-environment simulation loops across thousands of accelerators.
- Design rigorous experiments and failure analyses to isolate performance bottlenecks and communicate findings through clear write-ups.
- Maintain high code quality and architectural health across shared reinforcement learning and modeling codebases.
What they're looking for
- Bachelor's degree in Computer Science, Mathematics, Physics, a related quantitative field, or equivalent practical experience.
- 4 years of experience building, scaling, and debugging machine learning models using deep learning frameworks (e.g., JAX, PyTorch, or TensorFlow).
- Experience in one core area: Reinforcement Learning (RL), Post-Training (SFT/RLHF/RLAIF), Agentic Tool-Use, or Inference-Time Search.
Nice to have
- PhD in Computer Science, Machine Learning, Physics, or a related quantitative field.
- Experience training and managing models on large-scale distributed accelerator clusters (e.g., TPUs or GPUs).
- Experience designing asynchronous agent-environment simulation loops or large distributed post-training pipelines.
- Experience prototyping new hypotheses quickly while keeping shared codebases clean, robust, and production-grade.
Summarised by NextRaise from the employer’s description, which follows in full below.
Full description from employer
Minimum qualifications:
- Bachelor's degree in Computer Science, Mathematics, Physics, a related quantitative field, or equivalent practical experience.
- 4 years of experience building, scaling, and debugging machine learning models using deep learning frameworks (e.g., JAX, PyTorch, or TensorFlow).
- Experience in one core area: Reinforcement Learning (RL), Post-Training (SFT/RLHF/RLAIF), Agentic Tool-Use, or Inference-Time Search.
Preferred qualifications:
- PhD in Computer Science, Machine Learning, Physics, or a related quantitative field.
- Experience training and managing models on large-scale distributed accelerator clusters (e.g., TPUs or GPUs).
- Experience designing asynchronous agent-environment simulation loops or large distributed post-training pipelines.
- Experience prototyping new hypotheses quickly while keeping shared codebases clean, robust, and production-grade.
About the job
At DeepMind, the Planet-Scale Resources, Infrastructure and Systems Management (PRISM) team brings together researchers and engineers to advance the frontiers of AI reasoning and autonomous agentic systems. We reject the false tradeoff between research and execution, pursuing breakthroughs on open AI challenges while embedding directly into core teams to land those capabilities in production.
Artificial intelligence will be one of humanity’s most transformative inventions. At Google DeepMind, we are a pioneering AI lab with exceptional interdisciplinary teams focused on advancing AI development to solve complex global challenges and accelerate high-quality product innovation for billions of users. We use our technologies for widespread public benefit and scientific discovery, ensuring safety and ethics are always our highest priority.
US: $207000 - $300000 (USD) + 20% bonus target + equity + benefits
Learn more about benefits at Google.
Responsibilities
- Operate across the full research-and-engineering lifecycle of frontier reasoning and agentic systems.
- Work on unsolved problems in agentic reasoning, turning early exploratory prototypes into hardened production features for Gemini releases.
- Architect and optimize distributed post-training pipelines and agent-environment simulation loops across thousands of accelerators.
- Design rigorous experiments and failure analyses to isolate performance bottlenecks and communicate findings through clear write-ups.
- Maintain high code quality and architectural health across shared reinforcement learning and modeling codebases.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.