NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs
15 days ago
Apply with autofill
Apply with autofill
Kody·15 days ago
15 days ago

Senior Site Reliability Engineer

Hong Kong, Hong Kong, Hong KongSenior · 8+ yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at kody

How you compare FREE

?
Your scoreYour score: not yet known
→
67
Top 10%Top 10%: 67 out of 100

Top 10% of NextRaise users, across all roles in this function in Hong Kong.

Must-have skills for this role

  • aws
  • kubernetes
  • terraform
  • slo

PDF or DOCX · no account needed

Apply faster with autofill FREEkody uses Workable - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

What they're looking for

  • 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Job Summary

Kody is seeking a Senior Site Reliability Engineer (8+ years of experience) to drive the reliability, availability, scalability, and operational excellence of our global payment platform. Based in Hong Kong or Shenzhen, you will take end-to-end ownership of production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating across Europe, Asia, and North America.

Key Responsibilities

  • Incident Management & On-Call: Participate in a follow-the-sun production on-call rotation as a senior incident responder. Lead incident management during SEV1/SEV2 events to optimize MTTR and operational effectiveness.
  • Production Operations: Diagnose, triage, mitigate, and coordinate the resolution of complex production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
  • SLO & Reliability Engineering: Define, implement, and maintain SLOs, SLIs, error budgets, alerting standards, and operational readiness processes across distributed services.
  • Continuous Optimization: Drive systemic reliability improvements through infrastructure automation, observability enhancement, capacity planning, performance tuning, and post-incident root-cause analysis (RCA).
  • Security & Compliance: Partner with global engineering teams to strengthen architectural resilience, security posture, and operational maturity in PCI-DSS-regulated payment environments.
  • Technical Leadership: Mentor junior engineers, eliminate operational toil through automation, and influence engineering teams to adopt resilience-by-design practices.

Requirements

Qualifications & Requirements

  • Experience: 8+ years of hands-on experience in Site Reliability Engineering, Platform Engineering, DevOps, or Cloud Infrastructure roles supporting high-availability, mission-critical production systems.
  • Core Technical Stack: Strong expertise in AWS, Kubernetes (EKS), Terraform, PostgreSQL, Redis, Kafka, Linux, networking, and modern observability platforms (e.g., Datadog, Prometheus, Grafana).
  • Distributed Systems Mastery: Deep understanding of distributed systems architecture, high availability, disaster recovery, capacity planning, and microservices orchestration.
  • Domain Expertise: Proven track record operating in payment, banking, fintech, or other highly regulated environments with strict PCI-DSS, security, and uptime standards.
  • SRE Methodology: Deep knowledge of core SRE principles, including SLO/SLI design, error budget management, alert governance, and toil reduction.
  • Location & Communication: Based in Hong Kong or Shenzhen. Excellent command of English (written and spoken) to lead cross-functional incident responses and collaborate seamlessly with global teams.

Leadership & Operational Excellence

  • Ownership: Demonstrates strong end-to-end accountability for service reliability and customer impact under high pressure.
  • Structured Problem Solving: Applies a systematic and data-driven approach to troubleshooting, telemetry analysis, and incident resolution in complex distributed environments.
  • Crisis Management: Proven ability to command cross-functional incident response efforts, align stakeholders, and maintain clear communication during critical outages.
  • Engineering Culture: Champions a blameless post-incident culture, operational readiness, continuous learning, and technical mentorship.

Benefits

- Competitive Package

- A dynamic and innovative team

- Collaborative, inclusive working environment

Company

Kody
Hong Kong, Hong Kong, Hong Kong

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Kody's careers site·first seen 10 Sept 2026·last verified 10 Sept 2026·How we source jobs

Similar jobs

  • Senior Principal Site Reliability Engineer at BybitHong Kong SAR, Hong Kong–match not yet calculated
  • Sr. Manager, Site Reliability & Innovation, IT at citicclsaHong Kong–match not yet calculated
  • Site Reliability Engineer at quberesearchandtechnologiesHong kong–match not yet calculated
  • Production Engineer at quberesearchandtechnologiesHong Kong–match not yet calculated
  • Trading Systems Reliability Engineer (C++) at quberesearchandtechnologiesHong Kong–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in United States
  • Site Reliability Engineer jobs in India
  • Site Reliability Engineer jobs in United Kingdom
  • Retail Sales Associate jobs in United States