NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / AI / ML Researcher in United States of America
22 hours agoBe an early applicant
Apply with autofill
Apply with autofill
Designworkstalent·22 hours ago
22 hours agoBe an early applicant

Applied Researcher – Network Expert

Bellevue, United States of AmericaFull-timeHybridMid · 2-5 yearsAI / ML Researcher

Sign up free to see how well your resume matches this role.

Boost your chances at designworkstalent

How you compare FREE

?
Your scoreYour score: not yet known
→
49
Top 10%Top 10%: 49 out of 100

Top 10% of NextRaise users matched against AI / ML Researcher roles in United States.

Must-have skills for this role

  • infiniband
  • rocev2
  • nvlink
  • gpu networking

PDF or DOCX · no account needed

Apply faster with autofill FREEdesignworkstalent uses Ashby - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Hold the company's view of where AI networking is going.
  • Track vendor and hyperscaler roadmaps, research, standards work, and the startup and venture landscape across scale-up, scale-out and scale-across fabrics, topology and optics, collectives and the software above the fabric, operations and telemetry, and tenant-facing capability.
  • Formulate and validate the product and engineering thesis.
  • Own the company's fabric position: where the scale-up and scale-out boundary sits for our workloads, which transport we bet on and when, and what good telemetry and observability look like so fabric problems are diagnosable rather than inferred.
  • Help finance to formulate the numbers: cost per port, optics and cabling economics across pluggables and co-packaged options — reach, power draw, failure rates at scale — and the performance we can actually substantiate against what vendor benchmarks claim.
  • Own tenant-facing network capability for GPU-as-a-service: multi-tenant isolation and its performance cost, storage traffic alongside GPU traffic, and what we can commit to contractually — including an honest read of where we are undifferentiated against peers.
  • Make the work land commercially. Support sales and delivery in demanding customer conversations about cluster performance, feed product and go-to-market with what we can offer at what performance and price, and provide technical diligence on network vendors, partners, and prospective tuck-in targets.

What they're looking for

  • Deep professional experience in networking at scale
  • Hands-on experience with high-performance GPU or HPC fabrics at current generations
  • Demonstrated experience debugging real collective communication performance problems in production

Nice to have

  • Fluency with the landscape you would be scanning: the switch, NIC, and optics vendors, the standards bodies and consortia, and the startups attacking the fabric layer — and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.
  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources — vendor roadmaps, standards drafts, benchmark data, academic literature, your own testing — and producing a defensible position under genuine uncertainty. A PhD in a relevant technical field is one good route to this and is valued here; sustained industry research, standards-body work, or a body of internal technical assessments that changed real decisions are equally valid. Either way, the role turns on the second half: translating that work for engineering, product, go-to-market, and finance, because it informs all four.
  • Deep professional experience in computer networking and large-scale infrastructure.
  • Significant experience with GPU clusters, AI infrastructure, or high-performance computing environments.
  • Strong understanding of inter-GPU networking.
  • Hands-on experience with InfiniBand at current generations (NDR/XDR), high-performance Ethernet fabrics (RoCEv2, Spectrum-X, or Ultra Ethernet), or comparable HPC interconnects.
  • Practical understanding of NVLink / NVSwitch scale-up domains and how they interact with the scale-out fabric.
  • Experience debugging real collective communication performance problems — NCCL/RCCL, congestion, stragglers, topology mismatch — not just reading about them.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Applied Researcher – Network Expert

Location: Hybrid | Bellevue, WA (downtown)


About the Opportunity

Our client is seeking a Network Expert to join an Applied Research organization focused on the future of large-scale AI infrastructure.

This role is designed for a networking expert with deep experience in GPU-based computing environments and high-performance inter-GPU networking. As AI workloads increasingly rely on thousands of GPUs operating as a single logical system, the network connecting those GPUs becomes a critical component of overall infrastructure performance and scalability.

You will serve as a horizontal subject-matter expert, partnering with engineering and infrastructure leaders to evaluate technologies, architectures, and industry developments. You will help the organization understand where GPU networking is headed and what those developments mean for infrastructure, architecture, and investment decisions.

What You'll Do

  • Hold the company's view of where AI networking is going.

  • Track vendor and hyperscaler roadmaps, research, standards work, and the startup and venture landscape across scale-up, scale-out and scale-across fabrics, topology and optics, collectives and the software above the fabric, operations and telemetry, and tenant-facing capability. Right now that means questions like how fast Ultra Ethernet displaces the RoCEv2 fabrics most clusters actually run, where the scale-up domain should end now that NVLink has open challengers, what co-packaged optics does to power per port, and how to network a cluster that no longer fits in one building. Those specific questions will have changed within a year — holding the current version of them is the job.

  • Formulate and validate the product and engineering thesis. Turn that view into a defensible position on what we build, buy, or partner for in the fabric, pressure-tested against measured cluster performance, isolation requirements, and cost per port — and say so plainly when the evidence does not hold up.

  • Own the company's fabric position: where the scale-up and scale-out boundary sits for our workloads, which transport we bet on and when, and what good telemetry and observability look like so fabric problems are diagnosable rather than inferred. A fault that restarts a long training job is a direct cost, not an availability statistic.

  • Help finance to formulate the numbers: cost per port, optics and cabling economics across pluggables and co-packaged options — reach, power draw, failure rates at scale — and the performance we can actually substantiate against what vendor benchmarks claim.

  • Own tenant-facing network capability for GPU-as-a-service: multi-tenant isolation and its performance cost, storage traffic alongside GPU traffic, and what we can commit to contractually — including an honest read of where we are undifferentiated against peers.

  • Make the work land commercially. Support sales and delivery in demanding customer conversations about cluster performance, feed product and go-to-market with what we can offer at what performance and price, and provide technical diligence on network vendors, partners, and prospective tuck-in targets. Work closely with the Data Center Expert where the fabric meets the physical plant.


What We're Looking For

Required Qualifications

  • Deep professional experience in networking at scale,

  • Hands-on experience with high-performance GPU or HPC fabrics at current generations,

  • Demonstrated experience debugging real collective communication performance problems in production

Preferred Qualifications

  • Fluency with the landscape you would be scanning: the switch, NIC, and optics vendors, the standards bodies and consortia, and the startups attacking the fabric layer — and a view on which of them matter. Expect to be asked what you think is currently overhyped, and why.

  • Demonstrated ability to do research in the applied sense: taking an open question, investigating it from primary sources — vendor roadmaps, standards drafts, benchmark data, academic literature, your own testing — and producing a defensible position under genuine uncertainty. A PhD in a relevant technical field is one good route to this and is valued here; sustained industry research, standards-body work, or a body of internal technical assessments that changed real decisions are equally valid. Either way, the role turns on the second half: translating that work for engineering, product, go-to-market, and finance, because it informs all four.

  • Deep professional experience in computer networking and large-scale infrastructure.

  • Significant experience with GPU clusters, AI infrastructure, or high-performance computing environments.

  • Strong understanding of inter-GPU networking.

  • Hands-on experience with InfiniBand at current generations (NDR/XDR), high-performance Ethernet fabrics (RoCEv2, Spectrum-X, or Ultra Ethernet), or comparable HPC interconnects.

  • Practical understanding of NVLink / NVSwitch scale-up domains and how they interact with the scale-out fabric.

  • Experience debugging real collective communication performance problems — NCCL/RCCL, congestion, stragglers, topology mismatch — not just reading about them.

  • Understanding of multi-tenant network isolation and the security and performance tradeoffs involved in serving tenants on shared fabric.

  • Understanding of large-scale AI cluster configurations and the networking requirements associated with thousands of GPUs.

  • Ability to evaluate competing technologies and understand where the networking industry is heading.

  • Familiarity with NVIDIA and AMD GPU infrastructure ecosystems.

  • Strong analytical and technical communication skills.

  • Ability to operate as a horizontal technical expert and influence engineering decisions without necessarily owning implementation.

  • Ability to read technical papers and translate research concepts into practical engineering implications.

  • Ability to operate across engineering, research, and infrastructure organizations.

  • A track record of collaborating with researchers and engineers across groups and levels to shape long-term research directions and move research into practice.

  • Comfortable operating as an individual contributor with high ownership in a lean, early-stage team.

Location

  • Hybrid role based in downtown Bellevue, WA.

  • Approximately three days per week in the office.

  • Candidates elsewhere in the U.S. who are open to relocation are encouraged to apply.

  • U.S. work authorization is required. Visa sponsorship is not currently available.

  • Export control: this role involves technologies subject to U.S. export control regulations. Candidate eligibility may be subject to export control screening and, where applicable, licensing.

  • Travel: Willingness and ability to travel as needed internationally to data centers and co-locations (up to 25%)

 

Why Join?

  • High-impact technical role: Directly influence the technology direction of an organization building AI infrastructure at scale.

  • Ground-floor opportunity: Help establish technical strategy, architecture, processes, and culture within a growing organization.

  • High ownership: Operate as a senior individual contributor with substantial autonomy and direct access to senior technical leadership.

  • Cross-disciplinary exposure: Work across AI models, inference, accelerators, software systems, networking, infrastructure, and economics.

  • Cutting-edge technical problems: Work on multi-accelerator inference, intelligent routing, performance optimization, token economics, and compiler, kernel, and runtime technologies.

  • Research with practical impact: Turn emerging research and technology developments into decisions that directly affect engineering, product, commercial strategy, and investment.

  • Lean, senior environment: Work with a small group of highly experienced technical contributors rather than within a large management hierarchy.

Company

Designworkstalent
Bellevue, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Designworkstalent's careers site·first seen 19 Sept 2026·last verified 19 Sept 2026·How we source jobs

Similar jobs

  • Applied Researcher – Data Centers at designworkstalentBellevue, United States of America–match not yet calculated
  • Applied Researcher – AI Expert at designworkstalentBellevue, United States of America–match not yet calculated
  • Research Scientist, Networking Research - PhD New College Grad 2026 at NVIDIASanta Clara, United States of America–match not yet calculated
  • Researcher Senior (Health Services) at elevancehealthWILMINGTON, United States of America–match not yet calculated
  • Senior Researcher at MicrosoftRedmond, United States of America–match not yet calculated

Browse more jobs

  • AI / ML Researcher jobs in United States
  • Machine Learning Engineer jobs in United States
  • AI Engineer jobs in United States
  • Computer Vision Engineer jobs in United States
  • AI / ML Researcher jobs in United Kingdom
  • AI / ML Researcher jobs in Canada