NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Hardware Engineer in Canada
24 days ago
Apply with autofill
Apply with autofill
Huaweicanada·24 days ago
24 days ago

Senior Researcher - Edge AI Optimization/Hardware-Aware ML

Edmonton, CanadaFull-timeOn-siteMid · 2+ yearsHardware Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at huaweicanada

How you compare FREE

?
Your scoreYour score: not yet known
→
72
Top 10%Top 10%: 72 out of 100

Top 10% of NextRaise users matched against Hardware Engineer roles in Canada.

Must-have skills for this role

  • python
  • c++
  • pytorch
  • quantization

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Conduct research in hardware-aware neural network optimization (e.g., quantization-aware training, mixed precision, pruning, distillation, neural architecture search).
  • Develop novel approaches for latency/energy-aware training objectives and Pareto optimization (accuracy vs. compute vs. memory).
  • Prototype and evaluate techniques for efficient inference under device constraints (thermal limits, memory bandwidth, intermittent connectivity).
  • Publish and present findings internally and externally (papers, workshops, patents, technical blogs).
  • Optimize inference pipelines across pre/post-processing, scheduling, operator fusion, memory planning, and runtime execution.
  • Collaborate on or contribute to compilers / runtimes (e.g., TVM, MLIR, XLA, TensorRT, ONNX Runtime, TFLite, ExecuTorch) to improve operator coverage and performance.
  • Profile and optimize models with real device traces, addressing bottlenecks such as cache misses, memory bandwidth, kernel launch overhead, and CPU–NPU handoff.
  • Build and maintain hardware-aware benchmarking methodology and regression suites for edge targets (ARM CPU, mobile GPU, DSP, NPU).
  • Create deployment recipes for heterogeneous compute (CPU+GPU+NPU) including partitioning strategies and fallback paths.
  • Drive optimization for on-device personalization and incremental updates when needed (e.g., small adapters, efficient fine-tuning).
  • Partner with product engineering, platform teams, and hardware teams to translate device constraints into research targets and to transition research prototypes into production.
  • Mentor junior researchers/engineers, review experimental designs, and raise the quality bar for measurement rigor and reproducibility.

What they're looking for

  • PhD (or equivalent research experience) in Machine Learning, Computer Science, Electrical/Computer Engineering, or related field.
  • Strong programming skills in Python and C/C++ (or equivalent systems language).Experience building AI agent / harness / skill toolchains, including model evaluation, orchestration, and LLM-powered tooling. Hands-on experience with deep learning frameworks (e.g., PyTorch, TensorFlow, JAX) and deployment toolchains (e.g., ONNX, TFLite, TensorRT, TVM, MLIR-based stacks). Solid knowledge of performance profiling: latency measurement, memory profiling, kernel-level bottleneck analysis, and experimental rigor.
  • Proven publication record at top venues (e.g., NeurIPS/ICML/ICLR, MLSys, ASPLOS, ISCA, MICRO) and/or patents in ML efficiency.
  • 2+ years of relevant experience (research lab or industry) with demonstrated impact in at least one of:

Nice to have

  • Model compression (quantization/pruning/distillation)
  • Efficient architectures (MobileNet-like, MoE, efficient transformers, etc.)
  • ML systems/compilers/runtime optimization
  • hardware-aware optimization for edge deployment
  • Experience optimizing for specific edge hardware: ARM NEON, mobile GPUs, DSPs, NPUs, microcontrollers
  • Experience with distributed benchmarking, CI for performance regression, and reproducible experiment pipelines.
  • Understanding of power/thermal constraints and methodologies for measuring energy on device.
  • Experience with efficient LLM/VLM inference on edge (KV-cache optimization, quantized attention, speculative decoding, etc.).

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Huawei Canada has an immediate permanent opening for a Researcher.

About the team:

The Software-Hardware System Optimization Lab focuses on research and innovation in power efficiency and performance optimization for consumer devices. By leveraging the talents and capabilities of local academia and our team, we aim to build system-optimization capabilities for software and hardware across edge AI, multimedia, graphics, mobile gaming, and system software domains, thereby enhancing the user experience and performance competitiveness of Huawei's consumer device products.


About the job:

  • Conduct research in hardware-aware neural network optimization (e.g., quantization-aware training, mixed precision, pruning, distillation, neural architecture search).

  • Develop novel approaches for latency/energy-aware training objectives and Pareto optimization (accuracy vs. compute vs. memory).

  • Prototype and evaluate techniques for efficient inference under device constraints (thermal limits, memory bandwidth, intermittent connectivity).

  • Publish and present findings internally and externally (papers, workshops, patents, technical blogs).

  • Optimize inference pipelines across pre/post-processing, scheduling, operator fusion, memory planning, and runtime execution.

  • Collaborate on or contribute to compilers / runtimes (e.g., TVM, MLIR, XLA, TensorRT, ONNX Runtime, TFLite, ExecuTorch) to improve operator coverage and performance.

  • Profile and optimize models with real device traces, addressing bottlenecks such as cache misses, memory bandwidth, kernel launch overhead, and CPU–NPU handoff.

  • Build and maintain hardware-aware benchmarking methodology and regression suites for edge targets (ARM CPU, mobile GPU, DSP, NPU).

  • Create deployment recipes for heterogeneous compute (CPU+GPU+NPU) including partitioning strategies and fallback paths.

  • Drive optimization for on-device personalization and incremental updates when needed (e.g., small adapters, efficient fine-tuning).

  • Partner with product engineering, platform teams, and hardware teams to translate device constraints into research targets and to transition research prototypes into production.

  • Mentor junior researchers/engineers, review experimental designs, and raise the quality bar for measurement rigor and reproducibility.

  • Define technical roadmap areas (e.g., next-gen quantization, kernel optimization, model families for edge, compiler improvements).

About the ideal candidate:

  • PhD (or equivalent research experience) in Machine Learning, Computer Science, Electrical/Computer Engineering, or related field.

  • Strong programming skills in Python and C/C++ (or equivalent systems language).Experience building AI agent / harness / skill toolchains, including model evaluation, orchestration, and LLM-powered tooling. Hands-on experience with deep learning frameworks (e.g., PyTorch, TensorFlow, JAX) and deployment toolchains (e.g., ONNX, TFLite, TensorRT, TVM, MLIR-based stacks). Solid knowledge of performance profiling: latency measurement, memory profiling, kernel-level bottleneck analysis, and experimental rigor.

  • Proven publication record at top venues (e.g., NeurIPS/ICML/ICLR, MLSys, ASPLOS, ISCA, MICRO) and/or patents in ML efficiency.

  • 2+ years of relevant experience (research lab or industry) with demonstrated impact in at least one of:

    • Model compression (quantization/pruning/distillation)

    • Efficient architectures (MobileNet-like, MoE, efficient transformers, etc.)

    • ML systems/compilers/runtime optimization

    • hardware-aware optimization for edge deployment

  • Experience optimizing for specific edge hardware:

    • ARM NEON, mobile GPUs, DSPs, NPUs, microcontrollers

    • Experience with distributed benchmarking, CI for performance regression, and reproducible experiment pipelines.

    • Understanding of power/thermal constraints and methodologies for measuring energy on device.

    • Experience with efficient LLM/VLM inference on edge (KV-cache optimization, quantized attention, speculative decoding, etc.).

  • Technical Skills:

    • Quantization: PTQ/QAT, per-channel/per-tensor, calibration, smooth quant, GPTQ-like methods, mixed precision

    • Sparsity: structured pruning, N: M sparsity, hardware-friendly sparsity

    • Compiler techniques: graph rewriting, operator lowering, scheduling, kernel autotuning

    • Runtime techniques: memory arenas, tensor lifetime analysis, static vs dynamic shapes, batching strategies

    • Hardware fundamentals: cache hierarchy, SIMD, memory bandwidth, accelerator programming models

Additional Information:

Huawei Canada is committed to a fair, inclusive, and accessible recruitment process. If you require accommodation during any stage of the hiring process, please let us know and we will work with you to meet your needs.

All applications for this position are reviewed directly by our hiring team, we do not use artificial intelligence tools to screen or select candidates.

Company

Huaweicanada
Edmonton, Canada

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Huaweicanada's careers site·first seen 30 Aug 2026·last verified 15 Sept 2026·How we source jobs

Similar jobs

  • Hardware Developer Eng Co-op/Intern at NokiaCanada–match not yet calculated
  • NPI Hardware Co-op (8 month - January 2027) at cienaOttawa, Canada–match not yet calculated
  • Hardware Development Associate at keycafeVancouver, Canada–match not yet calculated
  • Hardware Electro-Optics Engineer at NokiaCanada–match not yet calculated
  • Junior Hardware Engineering Developer, FPGA at gdmsiOttawa, Canada–match not yet calculated

Browse more jobs

  • Hardware Engineer jobs in Canada
  • Electronics Engineer jobs in Canada
  • PCB Design Engineer jobs in Canada
  • Systems Integration Engineer jobs in Canada
  • Hardware Engineer jobs in United States
  • Hardware Engineer jobs in Germany