Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure
Sign up free to see how well your resume matches this role.
What you'll do
- Keep a massive, multi-cloud platform running for internal engineers
- Run incident response
- Provide hands-on support to internal teams
- Partner with developers to make services reliable at scale
- Manage services including Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms
- Operate across AWS, GCP, and on-premise Kubernetes
What they're looking for
- Experience managing massive, multi-cloud platforms (AWS, GCP)
- Experience with on-premise Kubernetes environments
- Experience with data and AI platform services (Spark, Flink, Airflow, Ray, Notebooks, LLM-based agent platforms)
- Experience with incident response
- Experience with infrastructure automation
Summarised by NextRaise from the employer’s description, which follows in full below.
Full description from employer
The Apple Services Engineering team (ASE) is one of the most exciting examples of Apple's long-held passion for combining art and technology. These are the people who power the App Store, Apple TV, Apple Music, Apple Podcasts, and Apple Books — at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.
Within ASE, the Apple Data Platform SRE team keeps a massive, multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure, automation, and customer success — running incident response, providing hands-on support to internal teams, and partnering with developers to make cutting-edge services like Spark, Flink, Airflow, Ray, Notebooks, and LLM-based agent platforms reliable at scale across AWS, GCP, and on-premise Kubernetes.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.