NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in Germany
1 month ago
Apply with autofill
Apply with autofill
1GLOBAL·Other·1 month ago
1 month ago

Senior Site Reliability Engineer (SRE)

Berlin, GermanyFull-timeMid · 2+ yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at 1GLOBAL

How you compare FREE

?
Your scoreYour score: not yet known
→
50
Top 10%Top 10%: 50 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in Germany.

Must-have skills for this role

  • linux
  • kubernetes
  • python
  • go

PDF or DOCX · no account needed

Apply faster with autofill FREE1GLOBAL uses Workable - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

About this role

About Us

1GLOBAL is a technology-driven global mobile communications provider helping enterprises and consumer brands deliver seamless connectivity worldwide. Powered by a best-in-class telecom platform, proprietary eSIM technology, and its own mobile core network across 15 countries, 1GLOBAL operates as a fully regulated telecommunications provider in 43 countries.


Founded in 2022, 1GLOBAL has rapidly become one of Europe's fastest-growing telecom technology companies, connecting more than 80 million people and devices globally. Our customers include leading banks, multinational enterprises, global retailers, travel companies, payment service providers, and digital-first businesses.


Headquartered in the Netherlands, with R&D hubs in Lisbon, Berlin, and São Paulo, 1GLOBAL employs over 550 professionals across 16 countries. With annual revenues exceeding US$200 million and a strong track record of profitability, we continue to invest in innovation, infrastructure, and international expansion as we redefine the future of global mobile connectivity.

About the Ideal Candidate

We are looking for a talented Senior Site Reliability Engineer (SRE) to join our Technology Department. We are open to hiring this role in Berlin, Germany. 

As a Senior SRE, you will be a senior individual contributor responsible for strengthening the stability, scalability, and reliability of our global infrastructure and services across both cloud and on-prem environments. You will work alongside SREs under the guidance of the SRE Team Lead, taking ownership of critical reliability domains and helping drive a data-driven reliability culture based on SLIs, SLOs, and error budgets.

Your mission will be to proactively identify weaknesses across systems and improve reliability through redundancy testing, automation, and observability. You will design, build, and operate the tools and processes that automatically detect, prevent, and recover from incidents, ensuring our services remain reliable and performant for customers around the world. 

This role collaborates closely with DevOps, Infrastructure, IP Network, and Security teams to maintain carrier-grade reliability standards across all layers of our platform.


About the Role

  • Act as a senior technical contributor within the SRE team, mentoring peers and setting the technical bar for reliability engineering. 
  • Define, measure, and maintain SLIs and SLOs for core infrastructure and customer-facing services. 
  • Plan and execute redundancy and resilience testing across service, infrastructure, and networking layers — validating failover, HA configurations, and disaster recovery readiness. 
  • Design and implement automated recovery mechanisms, self-healing workflows, and intelligent alerting systems. 
  • Drive incident response, root-cause analysis, and blameless post-mortems, and ensure implementation and tracking of corrective and preventive actions derived from them to achieve continuous improvement. 
  • Develop and enhance observability (metrics, logs, traces) using Prometheus, Grafana, Loki, and OpenTelemetry. 
  • Partner with Infrastructure and DevOps teams to ensure deployment safety, rollback policies, and configuration consistency. 
  • Proactively identify weaknesses through fault-injection, load, and chaos testing. 
  • Continuously reduce operational toil through automation and reliability tooling. 
  • Contribute to on-call practices, improving alert quality, runbooks, escalation procedures, and incident management processes. 
  • Perform capacity planning, performance benchmarking, and resilience audits across systems. 
  • Ensure compliance with security, reliability, and availability standards. 
  • Create and maintain internal documentation, playbooks, and operational guidelines for peers and users. 
  • Contribute to cloud cost-optimization initiatives, including reserved capacity planning, autoscaling design, storage tiering, workload right-sizing, and continuous anomaly detection.

Requirements

About You

Must-haves

  • A minimum of 5 years of experience in Site Reliability, Systems, or Infrastructure Engineering (including 2+ years in a dedicated SRE role). 
  • Strong expertise in Linux systems engineering, distributed systems, and networking. 
  • Proven experience building and running high-availability, mission-critical production systems. 
  • Hands-on experience with redundancy and failover testing, disaster recovery, and high-availability architecture validation. 
  • Deep understanding of monitoring, observability, and incident management principles.
  • Experience with Prometheus, Grafana, Loki, Thanos, and OpenTelemetry or similar tools. 
  • Proficiency in Python, Go, and Bash for automation and reliability tooling. 
  • Strong knowledge of Kubernetes, container orchestration, and service mesh architectures. 
  • Experience with AWS (EKS, EC2, VPC) and on-premises infrastructure integration. 
  • Proficiency in Infrastructure as Code tools such as Terraform. 
  • Understanding of networking fundamentals (routing, load balancing, BGP, DNS, VXLAN, etc.). 
  • Excellent analytical and problem-solving skills, capable of operating under pressure. 
  • Strong communication and collaboration skills across distributed and cross-functional teams.


Nice-to-haves

  • Experience in telecom, carrier-grade, or large-scale distributed systems environments.
  • Hands-on experience with chaos engineering and automated failure-scenario validation (e.g., simulating link or node failures).
  • Strong understanding of high-availability networking concepts.
  • Background in capacity planning, traffic engineering, and multi-region failover.
  • Experience building reliability dashboards and integrating SRE metrics into business KPIs or compliance reports.
  • Familiarity with security and resilience standards (ISO 27001, NIST SP 800-53).

Benefits

Why 1GLOBAL?

  • Growth Opportunities: Advance your career in one of the fastest growing telecommunications companies, expanding over 50% year-on-year under the leadership of successful tech entrepreneurs.
  • Major Transaction Exposure: Be in the driver’s seat for transactions that will have an impact on the future telco industry.
  • Work with a Talented Team: From the Board and the Founders to the Senior Management Team, you will collaborate daily with the most capable and renowned external advisors and constantly being exposed to talented and driven individuals.
  • Dynamic Work Environment: Thrive in a collaborative, fast-paced workplace where innovation is encouraged, and every contribution counts.
  • Professional Development: Work alongside industry experts to enhance your skills and knowledge in a cutting-edge field.
  • International Experience: Gain opportunities to work in different 1GLOBAL offices around the world as you grow within the company.
  • Open Communication Culture: Join a team where your ideas are heard, and open dialogue is encouraged, fostering a supportive and transparent work environment.
  • Get Things Done Attitude: Be part of a results-driven team that values efficiency, creativity, and the drive to make a tangible impact in the industry.


1GLOBAL is an equal opportunity employer, we value your character as much as your talent. Diversity drives our innovation, and we offer a collaborative, dynamic, and international work environment. We are excited for you to join our mission to revolutionise connectivity globally.

Other

Company

1GLOBALOther
Berlin, Germany

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from 1GLOBAL's careers site·first seen 17 Jun 2026·last verified 17 Sept 2026·How we source jobs

Similar jobs

  • Site Reliability Engineer -Cloudstore (m/f/d) at gridscaleKöln, Germany–match not yet calculated
  • Senior Platform & Reliability Engineer (all genders) at contaboRemote (Germany)–match not yet calculated
  • Site Reliability Engineer at trivagoDüsseldorf, Germany–match not yet calculated
  • System Reliability Engineer at sereactGermany–match not yet calculated
  • Site Reliability Engineer -Openstack (m/w/d) at gridscaleKöln, Germany–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in Germany
  • Systems Engineer jobs in Germany
  • DevOps Engineer jobs in Germany
  • Cloud Engineer jobs in Germany
  • Site Reliability Engineer jobs in United States
  • Site Reliability Engineer jobs in India