NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in United States of America
3 days ago
Apply with autofill
Apply with autofill
JPMorgan Chase·BFSI·3 days ago
3 days ago

Lead Site Reliability Engineer - Public Cloud Operational Excellence

Jersey City, United States of AmericaFull-timeMid · 5+ yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at JPMorgan Chase

How you compare FREE

?
Your scoreYour score: not yet known
→
44
Top 10%Top 10%: 44 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in United States.

Must-have skills for this role

  • sre
  • aws
  • slos
  • observability

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Drive consistency across SRE teams: establish and scale a “gold standard” for SLOs, on-call readiness, runbooks, postmortems, action tracking, and operational readiness across AWS/Azure/GCP.
  • Partner with Engineering and Product: embed reliability outcomes into roadmaps and delivery plans; convert incidents, support signals, and error budget trends into prioritized backlog with measurable impact.
  • Operational excellence & BPMs: design and run operating cadences (incident/stability reviews, KPI reviews, planning inputs), standardize intake/prioritization, and ensure closed-loop execution.
  • Risk governance: align reliability operations to risk/control expectations; operationalize recurring remediations (e.g., configuration drift, repeat findings) through centralized automation without sacrificing velocity.
  • Ticket reduction & automation opportunities: identify top drivers of support load and operational toil; build and maintain an automation opportunity pipeline; track adoption and deflection.
  • Platform stability metrics: own cross-platform reporting (SLO attainment, incident trends, MTTR/MTTI, change failure rate, ticket deflection, customer impact/cost of failure).
  • AI for reliability operations: apply AI/LLMs to improve triage, incident summarization, correlation/symptom mapping, and guardrailed automation; measure accuracy, safety, and outcomes.
  • Hands-on leadership: stay close to designs and critical implementations; lead systemic remediation and major incident response improvements.

What they're looking for

  • 5+ years in SRE / production engineering / platform reliability / infrastructure operations at enterprise scale
  • Demonstrated success driving cross-team standardization and measurable reliability outcomes through influence and operating mechanisms.
  • Deep knowledge of SLOs/SLIs, error budgets, observability, incident response, postmortems, and change reliability.
  • Strong experience partnering with Engineering and Product leadership to align priorities and deliver results.
  • Hands-on experience with AWS and/or Azure and automation/IaC fundamentals (e.g., Python/Go/Bash, Terraform, CI/CD).
  • Experience using AI/LLM-enabled approaches in operations (AIOps, AI-assisted troubleshooting, agentic workflows) with appropriate controls and measurement.

Nice to have

  • Improved SLO attainment and reduced Sev1/Sev2 frequency; fewer repeat incidents.
  • Reduced MTTR/MTTI and lower change failure rate; reduced customer minutes impacted (cost of failure).
  • Measurable ticket reduction/deflection via scaled automation and self-service adoption.
  • Consistent operating model adopted across SRE teams with trusted executive reporting.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

As a Lead Site Reliability Engineer at JPMorgan Chase within the Public Cloud team, you will blend hands-on engineering with program leadership to promote platform stability, ensure consistent execution across SRE teams, and partner closely with Engineering and Product to deliver measurable improvements in availability, support outcomes, and cost of failure.

Job Responsibilities

  • Drive consistency across SRE teams: establish and scale a “gold standard” for SLOs, on-call readiness, runbooks, postmortems, action tracking, and operational readiness across AWS/Azure/GCP.
  • Partner with Engineering and Product: embed reliability outcomes into roadmaps and delivery plans; convert incidents, support signals, and error budget trends into prioritized backlog with measurable impact.
  • Operational excellence & BPMs: design and run operating cadences (incident/stability reviews, KPI reviews, planning inputs), standardize intake/prioritization, and ensure closed-loop execution.
  • Risk governance: align reliability operations to risk/control expectations; operationalize recurring remediations (e.g., configuration drift, repeat findings) through centralized automation without sacrificing velocity.
  • Ticket reduction & automation opportunities: identify top drivers of support load and operational toil; build and maintain an automation opportunity pipeline; track adoption and deflection.
  • Platform stability metrics: own cross-platform reporting (SLO attainment, incident trends, MTTR/MTTI, change failure rate, ticket deflection, customer impact/cost of failure).
  • AI for reliability operations: apply AI/LLMs to improve triage, incident summarization, correlation/symptom mapping, and guardrailed automation; measure accuracy, safety, and outcomes.
  • Hands-on leadership: stay close to designs and critical implementations; lead systemic remediation and major incident response improvements.

Required qualifications, skills, and capabilities

  •  5+ years in SRE / production engineering / platform reliability / infrastructure operations at enterprise scale
  • Demonstrated success driving cross-team standardization and measurable reliability outcomes through influence and operating mechanisms.
  • Deep knowledge of SLOs/SLIs, error budgets, observability, incident response, postmortems, and change reliability.
  • Strong experience partnering with Engineering and Product leadership to align priorities and deliver results.
  • Hands-on experience with AWS and/or Azure  and automation/IaC fundamentals (e.g., Python/Go/Bash, Terraform, CI/CD).
  • Experience using AI/LLM-enabled approaches in operations (AIOps, AI-assisted troubleshooting, agentic workflows) with appropriate controls and measurement.

Preferred Qualifications

  • Improved SLO attainment and reduced Sev1/Sev2 frequency; fewer repeat incidents.
  • Reduced MTTR/MTTI and lower change failure rate; reduced customer minutes impacted (cost of failure).
  • Measurable ticket reduction/deflection via scaled automation and self-service adoption.
  • Consistent operating model adopted across SRE teams with trusted executive reporting.
BFSI

Company

JPMorgan ChaseBFSI
Jersey City, United States of America

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from JPMorgan Chase's careers site·first seen 17 Sept 2026·last verified 17 Sept 2026·How we source jobs

Similar jobs

  • Sr. Staff Site Reliability Engineer at earlywarningScottsdale, United States of America–match not yet calculated
  • Silicon Photonics Quality & Reliability Engineer at IntelAlbuquerque, United States of America–match not yet calculated
  • Site Reliability Engineer II at JPMorgan ChaseTampa, United States of America–match not yet calculated
  • Lead Site Reliability Engineer at JPMorgan ChaseOH, United States–match not yet calculated
  • Reliability Engineer at Finning InternationalElkford, United States of America–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in United States
  • Systems Engineer jobs in United States
  • Network Engineer jobs in United States
  • Platform Engineer jobs in United States
  • Site Reliability Engineer jobs in India
  • Site Reliability Engineer jobs in United Kingdom