NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Technical Consultant in United States of America
2 days ago
Apply with autofill
Apply with autofill
Designworkstalent·2 days ago
2 days ago

Technical Program Leader - AI Infrastructure

Bellevue, United States of AmericaFull-timeHybridSenior · 7+ yearsTechnical Consultant

Sign up free to see how well your resume matches this role.

Boost your chances at designworkstalent

How you compare FREE

?
Your scoreYour score: not yet known
→
38
Top 10%Top 10%: 38 out of 100

Top 10% of NextRaise users matched against Technical Consultant roles in United States.

PDF or DOCX · no account needed

Apply faster with autofill FREEdesignworkstalent uses Ashby - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Lead complex AI infrastructure and data center deployment programs from planning through production handoff.
  • Coordinate work across site readiness, racks, cabling, power, cooling, networking, storage, and GPU infrastructure.
  • Translate technical requirements into clear deployment plans, milestones, readiness criteria, and deliverables.
  • Partner closely with engineering, infrastructure, facilities, operations, supply chain, OEMs, contractors, and vendors.
  • Own program schedules, dependencies, critical paths, risks, issues, actions, and key decisions and RAID
  • Drive commissioning, integrated systems testing, burn-in, validation, punch-list resolution, and operational readiness.
  • Identify blockers early and work across teams to keep complex deployments moving.
  • Provide concise updates to technical and executive leadership on progress, risks, trade-offs, and decisions.
  • Help establish repeatable processes for deploying AI infrastructure efficiently and reliably at scale.

What they're looking for

  • 7+ years in data center infrastructure, technical deployment, mission-critical facilities, cloud, hyperscale, HPC, AI infrastructure, or a related environment.
  • Experience with several areas such as rack-and-stack, structured cabling, power, cooling, networking, storage, GPU infrastructure, commissioning, or operational readiness.
  • Experience leading technical workstreams and coordinating engineering teams, vendors, contractors, and other stakeholders.
  • Strong understanding of deployment planning, technical acceptance, defect management, change control, and production handoff.
  • Experience managing complex schedules, risks, dependencies, action items, and executive reporting.
  • Experience producing status reports, RAID logs, action trackers, readiness dashboards and leadership updates.
  • U.S. work authorization required; visa sponsorship is not currently available

Nice to have

  • GPU infrastructure: NVIDIA HGX, DGX, or other accelerator and rack-scale platforms.
  • High-performance networking: InfiniBand, RoCE, high-speed Ethernet, cluster fabrics, or large-scale network deployments.
  • AI/HPC storage: High-throughput or parallel storage supporting training and other demanding workloads.
  • High-density infrastructure: Direct liquid cooling, rear-door cooling, CDUs, high-density racks, and complex power requirements.
  • Cluster bring-up: Firmware, BIOS, drivers, provisioning, node diagnostics, burn-in, thermal/power testing, and RMA management.
  • Performance acceptance: Cluster benchmarking, communications testing, ML/HPC workloads, node stability, and production acceptance.
  • Production operations: Slurm, Kubernetes, GPU fleet management, monitoring, telemetry, and operational runbooks.
  • A degree in engineering, computer science, IT, data center operations, or a related technical field is preferred but not required.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Technical Program Leader – AI Infrastructure

We’re partnering with an early-stage technology company building and scaling advanced AI infrastructure—and we’re looking for a Technical Program Leader who wants to be close to the action.

This is not a traditional PMO role. You’ll be operating at the intersection of AI infrastructure, engineering, data center deployment, and large-scale technical execution, helping turn ambitious infrastructure plans into production-ready environments.

You’ll work with engineering teams, infrastructure leaders, vendors, contractors, and operations to deliver complex GPU and data center programs from site readiness through commissioning and production handoff.

For the right person, this is an opportunity to join at an important stage of growth and help establish the execution discipline, processes, and infrastructure that will support significant expansion.

What You’ll Do

  • Lead complex AI infrastructure and data center deployment programs from planning through production handoff.

  • Coordinate work across site readiness, racks, cabling, power, cooling, networking, storage, and GPU infrastructure.

  • Translate technical requirements into clear deployment plans, milestones, readiness criteria, and deliverables.

  • Partner closely with engineering, infrastructure, facilities, operations, supply chain, OEMs, contractors, and vendors.

  • Own program schedules, dependencies, critical paths, risks, issues, actions, and key decisions and RAID

  • Drive commissioning, integrated systems testing, burn-in, validation, punch-list resolution, and operational readiness.

  • Identify blockers early and work across teams to keep complex deployments moving.

  • Provide concise updates to technical and executive leadership on progress, risks, trade-offs, and decisions.

  • Help establish repeatable processes for deploying AI infrastructure efficiently and reliably at scale.

What We’re Looking For

You’re a technically credible program leader who understands infrastructure deployment firsthand. You don’t need to be the person configuring every server or switch, but you should be comfortable enough with the technology to challenge assumptions, understand dependencies, and earn the trust of engineering teams.

Required experience includes:

  • 7+ years in data center infrastructure, technical deployment, mission-critical facilities, cloud, hyperscale, HPC, AI infrastructure, or a related environment.

  • Experience with several areas such as rack-and-stack, structured cabling, power, cooling, networking, storage, GPU infrastructure, commissioning, or operational readiness.

  • Experience leading technical workstreams and coordinating engineering teams, vendors, contractors, and other stakeholders.

  • Strong understanding of deployment planning, technical acceptance, defect management, change control, and production handoff.

  • Experience managing complex schedules, risks, dependencies, action items, and executive reporting.

  • Experience producing status reports, RAID logs, action trackers, readiness dashboards and leadership updates.

AI / GPU / HPC Experience

Experience in one or more of these areas is particularly valuable:

  • GPU infrastructure: NVIDIA HGX, DGX, or other accelerator and rack-scale platforms.

  • High-performance networking: InfiniBand, RoCE, high-speed Ethernet, cluster fabrics, or large-scale network deployments.

  • AI/HPC storage: High-throughput or parallel storage supporting training and other demanding workloads.

  • High-density infrastructure: Direct liquid cooling, rear-door cooling, CDUs, high-density racks, and complex power requirements.

  • Cluster bring-up: Firmware, BIOS, drivers, provisioning, node diagnostics, burn-in, thermal/power testing, and RMA management.

  • Performance acceptance: Cluster benchmarking, communications testing, ML/HPC workloads, node stability, and production acceptance.

  • Production operations: Slurm, Kubernetes, GPU fleet management, monitoring, telemetry, and operational runbooks.

You do not need experience across every technology listed. We’re looking for someone with meaningful depth in several areas who can lead effectively across the broader deployment lifecycle.

The Scale

This is infrastructure built for modern AI workloads, not a conventional enterprise data center environment.

The programs may involve:

  • Multi-megawatt infrastructure and phased capacity deployments.

  • Hundreds or thousands of racks and significant GPU/accelerator capacity.

  • Multiple concurrent deployment workstreams or sites.

  • High-density rack designs, liquid cooling, high-speed networking, and complex power requirements.

  • Long-lead equipment, supply-chain constraints, and multiple technology vendors.

  • Cross-functional teams spanning engineering, facilities, supply chain, OEMs, contractors, and operations.

If you’ve delivered large-scale infrastructure and enjoy solving the problems that emerge when technology, facilities, vendors, and engineering all have to come together, this role will be highly relevant.

What Will Make You Successful

  • Technical credibility: You can engage directly with engineers and understand complex infrastructure decisions.

  • Execution mindset: You create structure without slowing teams down.

  • Hands-on leadership: You’re comfortable getting into the details when delivery is at risk and prioritize work onsite and keep delivering.

  • Strong communication: You can turn complex technical issues into clear actions and decisions.

  • Vendor accountability: You know how to manage scope, quality, milestones, and acceptance. You escalate issues or risks early and hold parties accountable.

  • End-to-end ownership: You stay engaged from initial planning through commissioning and production handoff.

  • Builder mentality: You enjoy working in an environment where processes are still evolving and you can help shape how things get done.

Education & Certifications

  • A degree in engineering, computer science, IT, data center operations, or a related technical field is preferred but not required.

  • Nice to have:

    • PMP, PRINCE2, ITIL, Uptime Institute, commissioning, data center, networking, cloud, GPU, or AI infrastructure certifications are a plus.

    • Related technology accreditations such as Cisco, Arista, NVIDIA, VMware, AWS, Azure or comparable infrastructure, network, cloud, GPU or AI platform certifications.

Work Arrangement

Bellevue, WA: Hybrid, with 3 days per week in the office for candidates within reasonable commuting distance.

Elsewhere in the U.S.: Fully remote for qualified candidates.

Travel: Up to 50% travel may be required

Work Eligibility: U.S. work authorization required; visa sponsorship is not currently available

Company

Designworkstalent
Bellevue, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Designworkstalent's careers site·first seen 18 Sept 2026·last verified 19 Sept 2026·How we source jobs

Similar jobs

  • Display Systems Engineering Co-op (Spring/Summer 2027) - Onsite at RTXcedar rapids, United States of America–match not yet calculated
  • Display Systems Engineering Co-Op (Summer/Fall 2027) - Onsite at RTXcedar rapids, United States of America–match not yet calculated
  • Member of Technical Staff - Engineering at patronusaiincSan Francisco, United States of America–match not yet calculated
  • Member of Technical Staff, Beneficial Deployments at openevidenceMiami, United States of America–match not yet calculated
  • Systems Engineering - Summer Intern 2027 - Onsite at globalhrCEDAR RAPIDS, United States of America–match not yet calculated

Browse more jobs

  • Technical Consultant jobs in United States
  • Functional Consultant jobs in United States
  • Solutions Consultant jobs in United States
  • Implementation Consultant jobs in United States
  • Technical Consultant jobs in India
  • Technical Consultant jobs in United Kingdom