Senior Manager, Core Infrastructure Engineering
Sign up free to see how well your resume matches this role.
About this role
Core Infrastructure Engineering within Oracle Cloud Infrastructure (OCI) is seeking an experienced Senior Manager, Site Reliability Engineering (SRE) to lead a team responsible for the reliability, availability, performance, and operational excellence of critical database and storage services that form the backbone of OCI.
This role combines people leadership with strong technical leadership. You will initially lead a team of approximately five experienced SREs and will be responsible for their development, prioritization of work, technical direction, and successful delivery. You will work closely with senior engineers and development teams to make architectural and operational decisions, resolve complex production challenges, and continuously improve the reliability and scalability of OCI services.
The ideal candidate has a strong background in Site Reliability Engineering, cloud infrastructure, storage, or database systems and has progressed into an engineering management role while maintaining strong technical depth. You should be comfortable leading experienced engineers, challenging technical assumptions, reviewing designs, and helping teams make sound engineering decisions in complex distributed environments.
You will play a key role in defining how the team operates, prioritizing reliability initiatives, managing operational risk, and ensuring that engineering investments address the most important reliability and customer-impacting problems. While this is a management role rather than a primarily hands-on engineering position, you will be expected to remain technically engaged and able to dive deeply into architecture, incidents, reliability challenges, and engineering trade-offs when needed
Qualifications
- 10+ years of relevant engineering experience, with a strong background in Site Reliability Engineering, cloud infrastructure, distributed systems, storage, database, or related infrastructure technologies.
- Proven experience managing and developing experienced engineers, ideally within SRE, infrastructure, platform, or cloud engineering teams.
- Strong technical foundation in operating large-scale, highly distributed production systems on cloud platforms such as OCI, AWS, GCP, or Azure.
-
Demonstrated ability to lead senior engineers and provide credible technical guidance.
- Strong understanding of reliability engineering practices, including availability, scalability, observability, incident management, capacity planning, performance management, and operational readiness.
- Experience troubleshooting complex production issues involving distributed systems, storage, databases, networking, or cloud infrastructure.
- Strong understanding of automation, infrastructure as code, CI/CD, configuration management, and modern cloud-native engineering practices.
- Working knowledge of technologies and tools such as Linux, Kubernetes, Docker, Terraform, Python, Go, Git, Jenkins, Grafana, or equivalent technologies.
- Experience operating business-critical, 24×7 high-availability production services.
- Ability to evaluate technical designs, identify operational risks, and guide engineering teams toward scalable and resilient solutions.
- Strong people leadership skills, including coaching, performance management, career development, prioritization, and team planning.
- Excellent communication and stakeholder management skills, with the ability to work effectively with senior engineers, engineering managers, and cross-functional teams.
- Strong organizational skills and the ability to operate independently in a complex, rapidly evolving engineering environment.
We are particularly interested in candidates who started their careers as SREs, infrastructure engineers, or engineers working on large-scale distributed systems and have subsequently moved into engineering management.
The successful candidate will combine the technical credibility required to lead experienced SREs with the people leadership skills needed to build and develop a high-performing team. They will be comfortable moving between people management, technical discussions, operational priorities, and broader engineering strategy.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.