NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / AI / ML Researcher in United States of America
21 hours agoBe an early applicant
Apply with autofill
Apply with autofill
Designworkstalent·21 hours ago
21 hours agoBe an early applicant

Applied Researcher – AI Expert

Bellevue, United States of AmericaFull-timeHybridMid · 2-5 yearsAI / ML Researcher

Sign up free to see how well your resume matches this role.

Boost your chances at designworkstalent

How you compare FREE

?
Your scoreYour score: not yet known
→
49
Top 10%Top 10%: 49 out of 100

Top 10% of NextRaise users matched against AI / ML Researcher roles in United States.

Must-have skills for this role

  • cuda
  • triton
  • ai models
  • machine learning

PDF or DOCX · no account needed

Apply faster with autofill FREEdesignworkstalent uses Ashby - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Hold the company's view of where data center infrastructure is going.
  • Track vendor and hyperscaler roadmaps, research, standards work, and the startup and venture landscape across power and energy, cooling and thermal, construction and delivery, rack and hall architecture, siting and regulation, and the economics that connect them.
  • Formulate and validate the product and engineering thesis.
  • Own the technical reference view of a client's hall at each GPU generation: density, power envelope, cooling topology, and what our operating and prospective sites can and cannot absorb.
  • Influence the numbers finance, pricing, and sales depend on: cost per rack, per MW, and per GPU-hour, and how they move with density, cooling approach, and silicon mix.
  • Make the work land commercially. Support sales and delivery in technically demanding customer and partner conversations, feed product and go-to-market with what we can credibly offer and when, and provide technical diligence on infrastructure partners, colocation providers, vendor reference designs, and prospective tuck-in targets.
  • Ideally also validate the power picture, since it is the constraint that governs everything else — interconnect availability and queue position, utility and PPA structures, on-site generation, long-lead equipment, and how all of it sets site selection and build sequencing.

What they're looking for

  • Significant hands-on experience with modern AI models,
  • Depth in inference rather than training alone,
  • Working fluency in one or more accelerator ecosystem, and
  • Hands-on depth at the compiler, kernel, or runtime layer (CUDA, Triton, ROCm/HIP, XLA, or similar).
  • Working fluency in more than one silicon ecosystem. CUDA plus ROCm, Cerebras, or another accelerator experience combined with a realistic view of what portability actually costs.

Nice to have

  • Fluency with the landscape you would be scanning: the frontier labs and open-weight model providers, the serving and inference startups, the silicon vendors, and the research groups doing the work that lasts — and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.
  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources — papers, model cards, vendor roadmaps, your own benchmarking — and producing a defensible position under genuine uncertainty. A PhD in a relevant field is one good route to this and is common among people with real depth here; sustained industry research, open-source contribution at depth, or a body of internal technical assessments that changed real decisions are equally valid. Combined with the ability to translate that work for engineering, product, go-to-market, and finance, because it informs all four.
  • Significant experience working with AI, machine learning, or AI model technologies.
  • Strong understanding of AI model architectures and how models are developed.
  • Ability to understand both the research and engineering implications of emerging AI technologies.
  • Experience working with one or more major model types, such as language, vision, audio, or multimodal models.
  • Hands-on depth in inference rather than only training — serving, optimization, and the practical work of getting latency and cost down without giving up quality.
  • Working knowledge of a modern serving stack (vLLM, SGLang, TensorRT-LLM, or equivalent) and of quantization, batching, and KV-cache techniques in production.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Applied Researcher – AI Expert

Location: Hybrid | Bellevue, WA (downtown)


About the Opportunity

Our client is seeking an experienced AI Expert / Applied Researcher to help shape how a fast-growing technology organization understands and applies the rapidly evolving landscape of AI models, architectures, inference technologies, and accelerator systems.

This role sits at the intersection of AI research, systems engineering, and infrastructure strategy. You will evaluate emerging technologies, translate research into practical engineering implications, and help guide decisions around AI infrastructure, inference optimization, model serving, accelerator platforms, and the economics of delivering AI workloads at scale.

This is a highly strategic individual contributor role with broad technical influence. You will work closely with engineering, product, infrastructure, finance, and commercial teams to help determine what technologies to build, adopt, partner for, or invest in.

 

What You'll Do

  • Hold the company's view of where data center infrastructure is going.

  • Track vendor and hyperscaler roadmaps, research, standards work, and the startup and venture landscape across power and energy, cooling and thermal, construction and delivery, rack and hall architecture, siting and regulation, and the economics that connect them. Right now that means questions like whether facility-level 800 VDC becomes the standard, how far liquid cooling has to go as racks move from roughly 200 kW toward a megawatt, how much of a build program can be moved into a factory, and whether behind-the-meter generation beats waiting in the interconnect queue. Those specific questions will have changed within a year — holding the current version of them is the job.

  • Formulate and validate the product and engineering thesis. Turn that view into a defensible position on what we build, buy, or partner for, pressure-tested against cost, schedule, and what the physical plant can actually support — and say so plainly when the evidence does not hold up.

  • Own the technical reference view of a client's hall at each GPU generation: density, power envelope, cooling topology, and what our operating and prospective sites can and cannot absorb. This includes getting on site to see them.

  • Influence the numbers finance, pricing, and sales depend on: cost per rack, per MW, and per GPU-hour, and how they move with density, cooling approach, and silicon mix. Be able to defend them under challenge.

  • Make the work land commercially. Support sales and delivery in technically demanding customer and partner conversations, feed product and go-to-market with what we can credibly offer and when, and provide technical diligence on infrastructure partners, colocation providers, vendor reference designs, and prospective tuck-in targets.

  • Ideally also validate the power picture, since it is the constraint that governs everything else — interconnect availability and queue position, utility and PPA structures, on-site generation, long-lead equipment, and how all of it sets site selection and build sequencing.


What We're Looking For

Required Qualifications

  • Significant hands-on experience with modern AI models,

  • Depth in inference rather than training alone,

  • Working fluency in one or more accelerator ecosystem, and

  • Hands-on depth at the compiler, kernel, or runtime layer (CUDA, Triton, ROCm/HIP, XLA, or similar).

  • Working fluency in more than one silicon ecosystem. CUDA plus ROCm, Cerebras, or another accelerator experience combined with a realistic view of what portability actually costs.

Preferred Qualifications

  • Fluency with the landscape you would be scanning: the frontier labs and open-weight model providers, the serving and inference startups, the silicon vendors, and the research groups doing the work that lasts — and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.

  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources — papers, model cards, vendor roadmaps, your own benchmarking — and producing a defensible position under genuine uncertainty. A PhD in a relevant field is one good route to this and is common among people with real depth here; sustained industry research, open-source contribution at depth, or a body of internal technical assessments that changed real decisions are equally valid. Combined with the ability to translate that work for engineering, product, go-to-market, and finance, because it informs all four.

  • Significant experience working with AI, machine learning, or AI model technologies.

  • Strong understanding of AI model architectures and how models are developed.

  • Ability to understand both the research and engineering implications of emerging AI technologies.

  • Experience working with one or more major model types, such as language, vision, audio, or multimodal models.

  • Hands-on depth in inference rather than only training — serving, optimization, and the practical work of getting latency and cost down without giving up quality.

  • Working knowledge of a modern serving stack (vLLM, SGLang, TensorRT-LLM, or equivalent) and of quantization, batching, and KV-cache techniques in production.

  • Ability to reason quantitatively about cost to serve, and to build models of it that hold up to scrutiny from finance and commercial teams.

Nice to Have Qualifications

  • Experience across multiple data center builds and multiple vendor reference designs.

  • Background in liquid cooling (direct-to-chip or immersion) or high-density rack deployments — meaningful above 100 kW, and increasingly relevant at the 200 kW-plus densities current GPU racks now draw.

  • Experience with commissioning, capacity planning, or handover of new data center halls into production.

  • Publications, patents, standards-body participation, or a visible external presence in the data center or infrastructure community.

  • Spent your education and career working deeply in data center architecture, but you have the foundational knowledge and intellectual curiosity to quickly understand new architectures and technologies as they emerge.


Location

  • Hybrid role based in downtown Bellevue, WA.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

  • Export control: this role involves technologies subject to U.S. export control regulations. Candidate eligibility may be subject to export control screening and, where applicable, licensing.

  • Travel: Willingness and ability to travel as needed internationally to data centers and co-locations (up to 25%)

 

Why Join?

  • High-impact technical role: Directly influence the technology direction of an organization building AI infrastructure at scale.

  • Ground-floor opportunity: Help establish technical strategy, architecture, processes, and culture within a growing organization.

  • High ownership: Operate as a senior individual contributor with substantial autonomy and direct access to senior technical leadership.

  • Cross-disciplinary exposure: Work across AI models, inference, accelerators, software systems, networking, infrastructure, and economics.

  • Cutting-edge technical problems: Work on multi-accelerator inference, intelligent routing, performance optimization, token economics, and compiler, kernel, and runtime technologies.

  • Research with practical impact: Turn emerging research and technology developments into decisions that directly affect engineering, product, commercial strategy, and investment.

  • Lean, senior environment: Work with a small group of highly experienced technical contributors rather than within a large management hierarchy.

Company

Designworkstalent
Bellevue, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Designworkstalent's careers site·first seen 19 Sept 2026·last verified 19 Sept 2026·How we source jobs

Similar jobs

  • Applied Researcher – Network Expert at designworkstalentBellevue, United States of America–match not yet calculated
  • Applied Researcher – Data Centers at designworkstalentBellevue, United States of America–match not yet calculated
  • Research Scientist, Networking Research - PhD New College Grad 2026 at NVIDIASanta Clara, United States of America–match not yet calculated
  • Researcher Senior (Health Services) at elevancehealthWILMINGTON, United States of America–match not yet calculated
  • Senior Researcher at MicrosoftRedmond, United States of America–match not yet calculated

Browse more jobs

  • AI / ML Researcher jobs in United States
  • Machine Learning Engineer jobs in United States
  • AI Engineer jobs in United States
  • Computer Vision Engineer jobs in United States
  • AI / ML Researcher jobs in United Kingdom
  • AI / ML Researcher jobs in Canada