NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Technical Consultant in United States of America
20 days ago
Apply with autofill
Apply with autofill
Hark·20 days ago
20 days ago

Technical Lead, On-Device AI Inference

San Jose, United States of AmericaFull-timeSenior · 8-12 yearsTechnical Consultant

Sign up free to see how well your resume matches this role.

Boost your chances at hark

How you compare FREE

?
Your scoreYour score: not yet known
→
38
Top 10%Top 10%: 38 out of 100

Top 10% of NextRaise users matched against Technical Consultant roles in United States.

Must-have skills for this role

  • gpus
  • npus
  • inference engines
  • high-performance computing

PDF or DOCX · no account needed

Apply faster with autofill FREEhark uses Greenhouse - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

About this role

About Hark

Hark is an artificial intelligence company building advanced, personalized intelligence. One that is proactive, multimodal, and capable of interacting with the world through speech, text, vision, and persistent memory.

We're pairing that intelligence with next-generation hardware to create a universal interface between humans and machines. While today's AI largely operates through chat boxes and decade-old devices, Hark is focused on what comes next: agentic systems that interact naturally with people and the real world.

To get there, we're developing multimodal models and next-generation AI hardware together - designed from the ground up as a single, unified interface for a new era of intelligent systems.

About the Role 

You'll own how Hark's models run on the silicon we ship: selecting the accelerators our devices are built around, co-designing architectures against real latency, memory, and power budgets, and building the low-level inference stack that turns a trained model into something that responds in milliseconds on a battery. You'll build and lead the team that does it. The ceiling on what our hardware can do is set here.

Responsibilities

  • Evaluate GPUs, NPUs, DSPs, and specialized accelerators for on-device deployment, and own the recommendation hardware decisions are made against.
  • Work with the foundation model and audio ML teams to shape architectures that meet deployment constraints before training locks them in.
  • Build the low-level execution layer, custom kernels, runtime systems, and compiler paths that transformer workloads run through on target hardware.
  • Partner with silicon vendors and internal hardware teams to bring up new accelerators and get efficient transformer execution on them early.
  • Hire and lead a team of engineers on performance-critical software, and set the technical bar for the inference stack.

Requirements

  • 8–12+ years in high-performance computing, including production workloads deployed on GPUs, NPUs, or specialized accelerators.
  • Deep understanding of attention, KV-cache behavior, quantization effects, and memory bandwidth limits.
  • You've designed or optimized inference engines, distributed runtimes, or ML compilers, and you write the kernels yourself when it matters.
  • Experience leading teams on performance-critical software. You've set direction on a stack, not just contributed to one.
  • You've taken a model from a research checkpoint to running on constrained hardware in a product people use.

Bonus Qualifications

  • Hands-on experience with Hexagon DSP, Ambiq-class MCUs, or comparable embedded AI silicon.
  • Experience with speech, audio, or streaming multimodal inference where latency is perceptible to the user.
  • Contributions to open-source inference or compiler toolchains (TensorRT, ONNX Runtime, TVM, MLIR, and similar).

Compensation

The US base salary range for this full-time position is between $300,000 - $500,000 annually.

The pay offered for this position may vary based on several individual factors, including job-related knowledge, skills, and experience. The total compensation package may also include additional components/benefits depending on the specific role. This information will be shared if an employment offer is extended.

Company

Hark
San Jose, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Hark's careers site·first seen 2 Sept 2026·last verified 8 Sept 2026·How we source jobs

Similar jobs

  • CM Systems Specialist at covestroBaytown, United States of America–match not yet calculated
  • Systems Specialists at bannerhealthBanner Health Corp Phoenix (2901 N Central Ave), United States of America–match not yet calculated
  • Systems Designer III (L5) at dtnaPortland, United States of America–match not yet calculated
  • Staff Engineer Systems at ngcSan Diego, United States of America–match not yet calculated
  • Technical Engineer (Production Liability Engineer) at M&T BankBuffalo, United States of America–match not yet calculated

Browse more jobs

  • Technical Consultant jobs in United States
  • Functional Consultant jobs in United States
  • Solutions Consultant jobs in United States
  • Implementation Consultant jobs in United States
  • Technical Consultant jobs in India
  • Technical Consultant jobs in United Kingdom