Senior Network HW System Engineer
Sign up free to see how well your resume matches this role.
About this role
Microsoft Silicon, Cloud Hardware, and Infrastructure Engineering (SCHIE) is the team behind Microsoft’s expanding Cloud Infrastructure and responsible for powering Microsoft’s “Intelligent Cloud” mission. SCHIE delivers the core infrastructure and foundational technologies for Microsoft's over 200 online businesses including Bing, MSN, Office 365, Xbox Live, Teams, OneDrive, and the Microsoft Azure platform globally with our server and data center infrastructure, security and compliance, operations, globalization, and manageability solutions. Our focus is on smart growth, high efficiency, and delivering a trusted experience to customers and partners worldwide and we are looking for passionate engineers to help achieve that mission.
Microsoft's Hardware Systems organization is developing AI-native silicon and system-level solutions to power the next generation of frontier AI models. The MAIA platform combines custom accelerators, advanced networking technologies, large-scale distributed infrastructure, and cloud-scale software to deliver industry-leading AI training and inference capabilities.
The Platform Systems Engineering (PSE) team is seeking a AI Network HW Systems Engineer to drive the architecture, bring-up, validation, optimization, and deployment of networking infrastructure for next-generation MAIA AI systems. This role sits at the intersection of networking hardware, systems architecture, AI infrastructure, and hyperscale deployment.
You will work across the entire networking stack, from high-speed SerDes interfaces, optics, cables, NICs, PHYs, and switch silicon to AI communication frameworks and distributed training workloads. You will collaborate with architects, silicon engineers, firmware developers, hardware designers, validation teams, OS developers, manufacturing partners, and Azure service teams to deliver scalable, reliable, and high-performance AI networking solutions.
This is a unique opportunity to influence the future of AI infrastructure while enabling Microsoft's next generation of frontier-scale AI systems.
#azure #MAIA #AI/ML #Networking Hardware
Responsibilities
AI Networking Architecture & Platform Development
Define and drive networking architectures for MAIA AI training and inference platforms, spanning scale-up and scale-out deployments. Partner across architecture, silicon, firmware, software, and Azure infrastructure teams to deliver networking solutions from concept through datacenter deployment while influencing future networking roadmaps.
Network Hardware Integration & System Bring-Up
Lead the integration, bring-up, and deployment of network hardware technologies including switches, NICs, PHYs, high-speed SerDes interfaces, optics, cables, and backplane solutions. Collaborate with internal teams, ODMs, and technology partners to ensure successful qualification and production readiness.
AI Fabric Validation, Performance & Reliability
Define and execute validation strategies for AI networking infrastructure, including functional, performance, scale, interoperability, reliability, and stress testing. Develop automated qualification frameworks and methodologies to ensure robust operation in rack-scale and cluster-scale AI environments.
Performance Optimization & Networking Efficiency
Analyze and optimize AI fabric performance across distributed training and inference workloads. Evaluate latency, bandwidth utilization, congestion behavior, and collective communication efficiency, translating workload requirements into scalable networking solutions and architecture recommendations.
High-Speed Interconnect & Emerging Technologies
Drive qualification and deployment of next-generation networking technologies, including high-speed copper and optical interconnects, PAM4-based SerDes, advanced optics, and future networking innovations such as LPO, LRO, CPO, and silicon photonics. Evaluate technology tradeoffs across performance, power, reliability, and scalability.
Debug, Telemetry & Automation
Lead root-cause analysis of networking and AI fabric issues spanning physical layer, network protocols, and distributed AI communication layers. Develop telemetry, diagnostics, automation, and fleet monitoring solutions that improve network reliability, accelerate issue resolution, and enhance engineering productivity.
Key Responsibilities
Drive end-to-end architecture, integration, validation, and deployment of networking infrastructure for MAIA AI systems across scale-up and scale-out environments.
Partner with silicon, firmware, software, hardware, and Azure infrastructure teams to define networking requirements and deliver scalable, reliable, and high-performance AI fabrics.
Lead bring-up, qualification, and optimization of network subsystems including switches, NICs, PHYs, optics, cables, and high-speed SerDes technologies.
Develop validation and performance methodologies for AI networking infrastructure, ensuring readiness across functionality, scale, reliability, and stress conditions.
Drive root-cause analysis, telemetry, and automation solutions to improve network resiliency, operational efficiency, and fleet health.
Evaluate and influence next-generation networking technologies and architectures required to support future AI workloads and hyperscale deployments.
Qualifications
Preferred Qualifications
AI Infrastructure Experience
Experience developing GPU, FPGA, TPU, AI accelerator, or HPC-based systems.
Familiarity with AI workload communication patterns and collective operations.
Understanding of large-scale distributed AI training environments.
Optical & Interconnect Expertise
Deep knowledge of:
Optical transceivers (DR4, DR8, FR4)
DSP architectures
TIAs and drivers
Optical link budgets
OMA, TDECQ, receiver sensitivity analysis
Silicon photonics technologies
Co-packaged optics architectures
Industry Standards
Experience with:
IEEE Ethernet standards
OIF specifications
CMIS management frameworks
Ultra Ethernet Consortium technologies
Future 224G and 448G ecosystems
Datacenter & Manufacturing
Experience with hyperscale datacenter deployments.
Knowledge of ODM, CM, and supplier ecosystems.
Experience managing products through EVT, DVT, PVT, and production ramps.
Software & Automation
Experience with Linux environments and automation frameworks.
Background developing telemetry, diagnostics, validation, and qualification tooling.
Familiarity with scripting and data analysis environments.
This position will be open for a minimum of 5 days, with applications accepted on an ongoing basis until the position is filled.
Microsoft is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to age, ancestry, citizenship, color, family or medical care leave, gender identity or expression, genetic information, immigration status, marital status, medical condition, national origin, physical or mental disability, political affiliation, protected veteran or military status, race, ethnicity, religion, sex (including pregnancy), sexual orientation, or any other characteristic protected by applicable local laws, regulations and ordinances. If you need assistance with religious accommodations and/or a reasonable accommodation due to a disability during the application process, read more about requesting accommodations.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.