Senior Site Reliability Engineer
Sign up free to see how well your resume matches this role.
What you'll do
- Support mission-critical cloud services and production operations
- Improve service reliability, reduce operational risk, automate repetitive tasks, and drive faster detection and resolution of issues
- Monitor service health, troubleshoot production issues, participate in incident response, improve observability, and implement reliability best practices
- Analyze recurring failures, build automation, support deployments, and contribute to capacity planning, disaster recovery, and operational readiness
- Work on different region/realm rollouts and deployments
- Forecast demands and respond to capacity needs
- Collaborate with software development teams to develop reliable and scalable infrastructures
- Perform data collection to maintain and optimize operations and reliability
- Leverage knowledge to perform incident response and/or maintenance tasks
- Provide health and performance reporting
- Identify opportunities for automation
- Communicate about services and identify and explain the potential impact of changes
Summarised by NextRaise from the employer’s description, which follows in full below.
Full description from employer
We are looking for a Site Reliability Engineer 3 to support mission-critical cloud services and production operations. The role focuses on improving service reliability, reducing operational risk, automating repetitive tasks, and driving faster detection and resolution of issues.
The engineer will work closely with development, infrastructure, security, and operations teams to monitor service health, troubleshoot production issues, participate in incident response, improve observability, and implement reliability best practices. This role also includes analyzing recurring failures, building automation, supporting deployments, and contributing to capacity planning, disaster recovery, and operational readiness.
Also works on number of different region/realm rollouts, deployments. Forecasts demands and responds to capacity needs. Collaborates with software development teams to develop reliable and scalable infrastructures. Performs data collection to maintain and optimize operations and reliability. Leverages knowledge to perform incident response and/or maintenance tasks. Provides health and performance reporting. Identifies opportunities for automation. Communicates about services and identifies and explains the potential impact of changes. Provides support for technology and document incidents. Experiments with new tools and assesses potential impact and develops knowledge of site reliability trends.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.