NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Learning & Development Specialist in United States of America
6 months ago
Apply with autofill
Apply with autofill
Hark·6 months ago
6 months ago

Infrastructure, Large-scale Training

San Jose, United States of AmericaFull-timeSenior · 5+ yearsLearning & Development Specialist

Sign up free to see how well your resume matches this role.

Boost your chances at hark

How you compare FREE

?
Your scoreYour score: not yet known
→
38
Top 10%Top 10%: 38 out of 100

Top 10% of NextRaise users matched against Learning & Development Specialist roles in United States.

Must-have skills for this role

  • infrastructure as code
  • ci/cd
  • rdma
  • infiniband

PDF or DOCX · no account needed

Apply faster with autofill FREEhark uses Greenhouse - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

About this role

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role

We are looking for a Member of Technical Staff, Infrastructure Compute to lead and manage large-scale GPU computing clusters powering our AI training and deployment workloads. You'll work at the intersection of systems engineering and machine learning infrastructure, owning the reliability, scalability, and efficiency of the compute platform that our research and engineering teams depend on. This is a high-impact, highly technical role suited for someone who thrives in complex distributed systems environments and cares deeply about infrastructure as a product.

Responsibilities

  • Design, implement, and maintain Infrastructure as Code (IaC) best practices to enable repeatable, auditable, and scalable cluster provisioning.
  • Enhance and harden CI/CD deployment pipelines to ensure robust, secure, and low-latency model service delivery across production environments.
  • Own and evolve stable training infrastructure operating at the scale of 10,000+ GPUs, including job scheduling, fault tolerance, and network fabric optimization.
  • Partner closely with ML researchers and engineers to understand compute bottlenecks and translate them into infrastructure improvements.
  • Monitor system health, define SLOs, and lead incident response for critical training and inference workloads.
  • Drive capacity planning, cost efficiency initiatives, and hardware lifecycle management across the GPU fleet.
  • Contribute to internal tooling and platform abstractions that improve developer experience for teams consuming compute resources.

Requirements

  • 5+ years of experience in infrastructure, systems, or platform engineering, with at least 2 years working in ML or HPC environments.
  • Demonstrated experience managing GPU clusters or large-scale distributed compute infrastructure.
  • Strong proficiency in at least one systems or infrastructure programming language.
  • Deep understanding of networking fundamentals (RDMA, InfiniBand, or RoCE a plus) relevant to high-throughput training workloads.
  • Experience with container orchestration, job scheduling, and multi-tenant resource management.
  • Proven track record owning production systems with high reliability requirements.
  • Strong debugging and observability skills across the full infrastructure stack.

Bonus Qualifications

  • Kubernetes (K8s) — particularly experience operating large, GPU-aware clusters.
  • Pulumi or similar modern IaC tooling.
  • Rust and/or Go for systems-level tooling and performance-critical services.
  • Familiarity with PyTorch and Ray for understanding workload patterns and integration requirements.

Compensation

The US base salary range for this full-time position is between $180,000 - $450,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components and benefits depending on the specific role. This information will be shared if an employment offer is extended.

Company

Hark
San Jose, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Hark's careers site·first seen 29 Aug 2026·last verified 8 Sept 2026·How we source jobs

Similar jobs

  • UAS Operations Pilot In Training at American Electric PowerAbilene, United States of America–match not yet calculated
  • TYSON FOODS -WILKESBORO FRESH PLANT: 1st Shift TRAINING CONE LINE General Labor PR01 at Tyson FoodsWilkesboro Plant, United States of America–match not yet calculated
  • Nuclear Training Leader at GE VernovaWilmington NC USA–match not yet calculated
  • Training Consultant - Mining at CaterpillarTucson, United States of America–match not yet calculated
  • Autism Aide-Training Included-Carlsbad, NM at csdautismservicesCarlsbad, United States of America–match not yet calculated

Browse more jobs

  • Learning & Development Specialist jobs in United States
  • Corporate Trainer jobs in United States
  • Culture & Engagement Specialist jobs in United States
  • Onboarding Specialist jobs in United States
  • Learning & Development Specialist jobs in United Kingdom
  • Learning & Development Specialist jobs in India