NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in United States of America
1 day agoBe an early applicant
Apply with autofill
Apply with autofill
Earlywarning·1 day ago
1 day agoBe an early applicant

Sr. Staff Site Reliability Engineer

Scottsdale, United States of AmericaSenior · 8-12 yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at earlywarning

How you compare FREE

?
Your scoreYour score: not yet known
→
44
Top 10%Top 10%: 44 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in United States.

Must-have skills for this role

  • devops
  • automation
  • observability
  • aws

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed, observed, operated, and recovered.
  • Use data, evidence, experimentation, and rigorous engineering analysis appropriate to the level to identify reliability risks, test assumptions, and guide technical decisions.
  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility.
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation.
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness.
  • Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices.
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle.
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning appropriate to the level.
  • Provides senior technical leadership and escalation support for significant production incidents and drives improvements to incident response and sustainable on-call practices across a pillar or broad technical domain.
  • Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices.

What they're looking for

  • Typically 12+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline.
  • Experience with software development or scripting using one or more modern programming languages.
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level.
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level.
  • Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role.

Nice to have

  • Hands-on experience with AWS is preferred, or comparable experience with another major cloud platform such as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI).
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems.
  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation.
  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness appropriate to the level.
  • Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness.
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

At Early Warning, we’ve powered and protected the U.S. financial system for over thirty years with cutting-edge solutions like Zelle®, Paze℠, and so much more. As a trusted name in payments, we partner with thousands of institutions to increase access to financial services and protect transactions for hundreds of millions of consumers and small businesses.

Positions located in Scottsdale, San Francisco, Chicago, or New York follow a hybrid work model to allow for a more collaborative working environment.

Candidates responding to this posting must independently possess the eligibility to work in the United States, for any employer, at the date of hire. This position is ineligible for employment Visa sponsorship.

The Senior Staff Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle.

The role provides broad technical leadership across a major technology domain, pillar, or portfolio and aligns multiple teams around sustainable reliability practices and technical direction.

Role Summary 

The Senior Staff Site Reliability Engineer applies software engineering and systems engineering practices to improve the reliability, resilience, scalability, and operational health of production services. The role partners with Software Engineering and other technology teams to ensure reliability, observability, recoverability, performance, and operational readiness are engineered into systems throughout their lifecycle. 

The role provides broad technical leadership across a major technology domain, pillar, or portfolio and aligns multiple teams around sustainable reliability practices and technical direction. 

Core Responsibilities 

  • Use software engineering, automation, and DevOps principles and practices to continually improve how services are built, tested, deployed, observed, operated, and recovered. 
  • Use data, evidence, experimentation, and rigorous engineering analysis appropriate to the level to identify reliability risks, test assumptions, and guide technical decisions. 
  • Define, implement, or improve SLIs, SLOs, error budgets, and other service-health measures appropriate to the scope of responsibility. 
  • Improve observability through metrics, logging, tracing, monitoring, alerting, dashboards, and service-health instrumentation. 
  • Drive continuous improvement across CI/CD, observability, deployment practices, Infrastructure as Code, automation, testing, incident response, capacity management, resilience, and operational readiness. 
  • Identify recurring or systemic production issues and translate operational experience into improvements in code, architecture, automation, tooling, and engineering practices. 
  • Partner with Software Engineering teams to incorporate reliability, resiliency, scalability, performance, observability, recoverability, and operational readiness throughout the development lifecycle. 
  • Participate in or lead incident response, troubleshooting, service restoration, and blameless post-incident learning appropriate to the level. 
  • Provides senior technical leadership and escalation support for significant production incidents and drives improvements to incident response and sustainable on-call practices across a pillar or broad technical domain. 
  • Reduce operational toil and unnecessary manual intervention through software, automation, reusable patterns, and better engineering practices. 

Leveling Intent 

Senior Staff represents the transition from cross-team technical leadership to broad technical direction. Senior Staff SREs align teams around common reliability objectives and multiply capability across a major domain or pillar while creating sustainable technical leadership depth. 

Level Expectations 

  • Acts as a force multiplier across a major domain or pillar by aligning engineering approaches, developing technical leaders, spreading knowledge and reusable solutions, and reducing organizational dependency on individual expertise. 
  • Demonstrates software engineering, systems thinking, troubleshooting, and production reliability capabilities appropriate to the level. 
  • Applies evidence-driven reasoning and technical rigor to distinguish observed facts from assumptions and make defensible engineering recommendations. 
  • Shares knowledge and contributes to sustainable engineering capability rather than creating dependency on individual expertise. 
  • Establishes technical direction across a major domain, pillar, or portfolio. 
  • Develops Staff and Senior engineers and creates sustainable technical leadership depth. 

Minimum Qualifications 

  • Typically 12+ years of relevant professional experience in Software Engineering, Site Reliability Engineering, Systems Engineering, Cloud/Platform Engineering, DevOps, Infrastructure Engineering, Architecture where applicable, or a comparable technical discipline. 
  • Experience with software development or scripting using one or more modern programming languages. 
  • Experience with software engineering principles, distributed systems, production troubleshooting, automation, and observability appropriate to the level. 
  • Experience with public cloud technologies and architectures, preferably AWS, along with infrastructure, networking, Linux/Unix, and modern application architectures appropriate to the level. 
  • Demonstrated analytical, problem-solving, communication, and collaboration skills appropriate to the scope of the role. 

Preferred Qualifications 

  • Hands-on experience with AWS is preferred, or comparable experience with another major cloud platform such as Microsoft Azure, Google Cloud Platform (GCP), or Oracle Cloud Infrastructure (OCI). 
  • Experience developing, deploying, operating, or improving highly available production software or distributed systems. 
  • Experience with CI/CD, Infrastructure as Code, containers or orchestration, observability, monitoring, alerting, and software-delivery automation. 
  • Experience with SLIs, SLOs, error budgets, incident management, performance analysis, capacity management, resilience testing, disaster recovery, or operational readiness appropriate to the level. 
  • Experience creating reusable automation, tooling, platforms, patterns, or practices that improve engineering effectiveness. 
  • Bachelor's degree in Computer Science, Software Engineering, Computer Engineering, Information Systems, or a related technical field, or equivalent practical experience.  

The base pay scale for this position in:
Phoenix, AZ/ Chicago, IL in USD per year is: $150,000 - $200,000.
Additionally, candidates are eligible for a discretionary incentive plan and benefits.

This pay scale is subject to change and is not necessarily reflective of actual compensation that may be earned, nor a promise of any specific pay for any specific candidate, which is always dependent on legitimate factors considered at the time of job offer. Early Warning Services takes into consideration a variety of factors when determining a competitive salary offer, including, but not limited to, the job scope, market rates and geographic location of a position, candidate’s education, experience, training, and specialized skills or certification(s) in relation to the job requirements and compared with internal equity (peers). The business actively supports and reviews wage equity to ensure that pay decisions are not based on gender, race, national origin, or any other protected classes.

Some of the Ways We Prioritize Your Health and Happiness 

  •  Healthcare Coverage – Competitive medical (PPO/HDHP), dental, and vision plans as well as company contributions to your Health Savings Account (HSA) or pre-tax savings through flexible spending accounts (FSA) for commuting, health & dependent care expenses.

  • 401(k) Retirement Plan – Featuring a 100% Company Safe Harbor Match on your first 6% deferral immediately upon eligibility.

  • Paid Time Off – Flexible Time Off for Exempt (salaried) employees, as well as generous PTO for Non-Exempt (hourly) employees, plus 11 paid company holidays and a paid volunteer day.

  • 12 weeks of Paid Parental Leave

  • Maven Family Planning – provides support through your Parenting journey including egg freezing, fertility, adoption, surrogacy, pregnancy, postpartum, early pediatrics, and returning to work.

 

And SO much more! We continue to enhance our program, so be sure to check our Benefits page here for the latest. Our team can share more during the interview process!

 

Early Warning Services, LLC (“Early Warning”) considers for employment, hires, retains and promotes qualified candidates on the basis of ability, potential, and valid qualifications without regard to race, religious creed, religion, color, sex, sexual orientation, genetic information, gender, gender identity, gender expression, age, national origin, ancestry, citizenship, protected veteran or disability status or any factor prohibited by law, and as such affirms in policy and practice to support and promote equal employment opportunity and affirmative action, in accordance with all applicable federal, state, and municipal laws. The company also prohibits discrimination on other bases such as medical condition, marital status or any other factor that is irrelevant to the performance of our employees. 

Early Warning Services LLC is a proud participant in E-Verify, a federal program to help ensure a legal and authorized workforce. As part of our hiring process, we electronically verify the employment eligibility of all new hires through E-Verify. For more information on your rights and responsibilities under E-Verify please visit Home | E-Verify.

Company

Earlywarning
Scottsdale, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Earlywarning's careers site·first seen 19 Sept 2026·last verified 19 Sept 2026·How we source jobs

Similar jobs

  • Silicon Photonics Quality & Reliability Engineer at IntelAlbuquerque, United States of America–match not yet calculated
  • Site Reliability Engineer II at JPMorgan ChaseTampa, United States of America–match not yet calculated
  • Lead Site Reliability Engineer at JPMorgan ChaseOH, United States–match not yet calculated
  • Reliability Engineer at Finning InternationalElkford, United States of America–match not yet calculated
  • Senior Reliability Engineer, Technical Lead at muonspaceSan Jose, United States of America–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in United States
  • Systems Engineer jobs in United States
  • Network Engineer jobs in United States
  • Platform Engineer jobs in United States
  • Site Reliability Engineer jobs in India
  • Site Reliability Engineer jobs in United Kingdom