NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs
1 month ago
Apply with autofill
Apply with autofill
NVIDIA·Semiconductors·1 month ago
1 month ago

Manager, Performance Research and Analysis

Yokneam, IsraelFull-timeHybridSenior · 5-10 yearsPerformance Marketing Manager

Sign up free to see how well your resume matches this role.

Boost your chances at NVIDIA

How you compare FREE

?
Your scoreYour score: not yet known
→
35
Top 10%Top 10%: 35 out of 100

Top 10% of NextRaise users, across all roles in this function in Israel.

Must-have skills for this role

  • rdma
  • nccl
  • gpu
  • networking

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation NVIDIA AI GPU clusters, focusing on large-scale distributed training and inference workloads.
  • Deeply evaluate and optimize NVIDIA Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.
  • Work on performance research and analysis of NVIDIA DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.
  • Drive the strategy for performance observability and dashboards across next-generation NVIDIA data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.
  • Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.

What they're looking for

  • B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.
  • 8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.
  • 3+ years of experience as an engineering team manager leading technical performance or R&D teams.
  • Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).
  • Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.
  • Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.

Nice to have

  • Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.
  • Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.
  • Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.
  • Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

NVIDIA is seeking a highly skilled and versatile Performance Research and Analysis Manager to join our Performance Group. This role will drive end-to-end performance strategy and execution for next-generation NVIDIA data centers and solutions based on GPU systems, NIC, Switch, DPU and Networking technologies. The ideal candidate will oversee, evaluating, and optimizing end-to-end AI GPU cluster-level performance for scaling out large scale distributed training and inference jobs communication. The role will focus heavily on RDMA, Networking Protocols, Collective Communication, Congestion Control, and Load Balancing algorithms. Secondarily, you will lead NVIDIA DPUs and Storage technologies for N-S use cases to support AI Inference jobs. Third, you will drive our Performance Dashboards and Observability for cluster-level performance analysis from a stream line telemetry across NICs, Switches, GPUs, and NVlink.

What you'll be doing:

  • Drive end-to-end performance strategy, characterization, test plans, and optimization for next-generation NVIDIA AI GPU clusters, focusing on large-scale distributed training and inference workloads.

  • Deeply evaluate and optimize NVIDIA Networking core technologies performance, including RDMA/PRDMA, networking protocols, collective communication (NCCL), congestion control, and load-balancing algorithms.

  • Work on performance research and analysis of NVIDIA DPUs and storage technologies in North-South (N-S) use cases and deployment scenarios to maximize performance and efficiency for AI inference jobs.

  • Drive the strategy for performance observability and dashboards across next-generation NVIDIA data center solutions and supercomputers by leveraging scalable, streamlined telemetry pipelines to build performance dashboards and automated analytics based on real-time performance metrics across NICs, Switches, GPUs, and NVLink boundaries.

  • Perform deep root-cause analysis (RCA) on complex multi-node performance bottlenecks, driving actionable mitigation plans across hardware, firmware, and software teams.

What we need to see:

  • B.Sc. or M.Sc. in Computer Science, Computer Engineering, Software Engineering, or equivalent technical experience.

  • 8+ overall years of experience and deep expertise in High Performance Networking, RDMA, and Systems level performance.

  • 3+ years of experience as an engineering team manager leading technical performance or R&D teams.

  • Hands-on experience analyzing and optimizing collective communication (e.g., NCCL, MPI) and network traffic patterns for large-scale distributed AI workloads (LLM training and inference).

  • Hands-on experience designing, deploying, and customizing Grafana dashboards for cluster monitoring, alerting, and data visualization.

  • Exceptional cross-team leadership, analytical thinking, and communication skills to drive alignment across hardware, software, and architecture groups.

Ways to stand out from the crowd:

  • Proven track record of optimizing NCCL, RDMA/RoCEv2, and custom collective algorithms specifically tailored for multi-thousand GPU deployments running LLMs or Mixture-of-Experts (MoE) architectures.

  • Deep experience tuning advanced network traffic mechanisms such as adaptive routing, PFC/ECN congestion control, and packet-spraying technologies.

  • Experience building autonomous performance-driven tools, AI-assisted root cause analysis agents, or automated regression frameworks for continuous cluster-level performance evaluation.

  • Hands-on experience developing custom Grafana plugins, complex dashboard panels, or integrated alert management workflows using PromQL/LogQL for hyperscale or HPC environments.

NVIDIA is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our current and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national origin, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law.

#LI-Hybrid

Semiconductors

Company

NVIDIASemiconductors
Yokneam, Israel

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from NVIDIA's careers site·first seen 31 Jul 2026·last verified 8 Sept 2026·How we source jobs

Similar jobs

  • PPC Manager at humanzRamat Gan, Israel–match not yet calculated
  • B2B Performance Marketing Manager at healtheeTel Aviv, Israel–match not yet calculated
  • Performance Marketing Manager (Web) at imagineartIsrael–match not yet calculated
  • Performance Manager / Media Buyer at techbizglobalRishon Lezion, Israel–match not yet calculated
  • Senior Manager, DPU Performance and System Validation at nvidiaYokneam, Israel–match not yet calculated

Browse more jobs

  • Performance Marketing Manager jobs in United States
  • Performance Marketing Manager jobs in United Kingdom
  • Performance Marketing Manager jobs in Germany
  • Retail Sales Associate jobs in United States