NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Software Engineer in United States of America
10 hours agoBe an early applicant
Apply with autofill
Apply with autofill
Riverai·10 hours ago
10 hours agoBe an early applicant

Software Engineer, GPU Kernels

Palo Alto, United States of AmericaMid · 2-5 yearsSoftware Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at riverai

How you compare FREE

?
Your scoreYour score: not yet known
→
65
Top 10%Top 10%: 65 out of 100

Top 10% of NextRaise users matched against Software Engineer roles in United States.

Must-have skills for this role

  • cuda
  • triton
  • cutlass
  • cute

PDF or DOCX · no account needed

Apply faster with autofill FREEriverai uses Greenhouse - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
  • Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.
  • Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
  • Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
  • Profile real workloads and integrate improvements into training and inference runtimes.
  • Build reproducible benchmarks that verify correctness, gradients, and performance.

What they're looking for

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
  • Proficiency in C++ and Python.
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
  • Strong debugging and profiling skills, with a collaborative approach to engineering.

Nice to have

  • Experience optimizing for NVIDIA Blackwell or Hopper GPUs.
  • Work on attention, mixture-of-experts kernels, grouped GEMMs, or low-rank adapters.
  • Experience implementing backward passes and validating gradients.
  • Familiarity with FP8, FP4, and quantized weight layouts.
  • Experience integrating custom operators into PyTorch, SGLang, vLLM, or similar frameworks.
  • Open-source contributions or a track record of shipping substantial kernel optimizations.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

At River AI, our mission is to create personal AI owned and shaped by each individual. To achieve this, we are rewriting the entire stack from scratch: personal hardware for local inference, bespoke training infrastructure, next-generation UIs, and frontier deep learning research.

Who we are

We are scientists, engineers, and builders from the industry's top tech companies and AI labs. We bring a proven track record of scaling consumer systems for hundreds of millions of users and architecting the pre-training infrastructure behind today's frontier models.

About the Role

We are looking for exceptional GPU kernel engineers to build the compute primitives behind River’s training and inference infrastructure. Your goal is to make large models faster to train and more efficient to serve.

You will own performance-critical operations, including attention, matrix multiplication, mixture-of-experts execution, and low-precision computation. Working closely with researchers and systems engineers, you will identify bottlenecks, implement kernels, validate correctness, and bring improvements into production.

What You’ll Do

  • Build fast GPU kernels for attention, matrix multiplication, expert routing, and related operations.
  • Optimize memory access, tiling, and synchronization to make efficient use of GPU hardware.
  • Develop FP8, FP4, and mixed-precision kernels while preserving numerical correctness.
  • Accelerate fine-tuning and RL through optimized adapters, backward passes, and fused operations.
  • Profile real workloads and integrate improvements into training and inference runtimes.
  • Build reproducible benchmarks that verify correctness, gradients, and performance.

Skills & Qualifications

Minimum Qualifications:

  • Bachelor’s degree in Computer Science, Computer Engineering, or equivalent practical experience.
  • Experience optimizing GPU kernels with CUDA, Triton, CUTLASS, CuTe, or comparable tools.
  • Strong understanding of GPU architecture, memory hierarchies, and parallel execution.
  • Proficiency in C++ and Python.
  • Strong foundations in linear algebra, floating-point arithmetic, and numerical computing.
  • Strong debugging and profiling skills, with a collaborative approach to engineering.

Preferred Qualifications: (We encourage you to apply even if you don't meet all of these)

  • Experience optimizing for NVIDIA Blackwell or Hopper GPUs.
  • Work on attention, mixture-of-experts kernels, grouped GEMMs, or low-rank adapters.
  • Experience implementing backward passes and validating gradients.
  • Familiarity with FP8, FP4, and quantized weight layouts.
  • Experience integrating custom operators into PyTorch, SGLang, vLLM, or similar frameworks.
  • Open-source contributions or a track record of shipping substantial kernel optimizations.

Logistics & Benefits

  • Location: Palo Alto, California.
  • Compensation: Depending on experience and skills the expected base pay is $200,000 - $420,000 USD per year.
  • Benefits: Comprehensive health, dental, and vision insurance; unlimited PTO; and relocation assistance as needed.
  • Visa Sponsorship: We sponsor visas and are committed to supporting the process for the right candidate.

Company

Riverai
Palo Alto, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from riverai's careers site·first seen 12 Sept 2026·last verified 12 Sept 2026·How we source jobs

Similar jobs

  • Software Engineer II - LoopNet - Irvine, CA at costarIrvine (US), United States of America–match not yet calculated
  • Senior Software Engineer at brooksautoFremont, United States of America–match not yet calculated
  • Senior Systems Programmer Analyst at bbhJersey City, United States of America–match not yet calculated
  • Software Engineer, Distributed Training at riveraiPalo Alto, United States of America–match not yet calculated
  • Software Engineer, Inference Systems at riveraiPalo Alto, United States of America–match not yet calculated

Browse more jobs

  • Software Engineer jobs in United States
  • Backend Engineer jobs in United States
  • Full Stack Engineer jobs in United States
  • API Engineer jobs in United States
  • Software Engineer jobs in India
  • Software Engineer jobs in United Kingdom