NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in Germany
13 days ago
Apply with autofill
Apply with autofill
Zalando·E-commerce·13 days ago
13 days ago

Senior Monitoring & Observability Engineer (all genders)

Ansbach, GermanySenior · 5-8 yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at Zalando

How you compare FREE

?
Your scoreYour score: not yet known
→
50
Top 10%Top 10%: 50 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in Germany.

Must-have skills for this role

  • prometheus
  • grafana
  • aws
  • new relic

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Monitoring, Alerting & Observability: Implementing tools and dashboards to track system health and business processes in real-time, ensuring the team is proactively notified of any technical or functional anomalies.
  • Incident & Problem Management (L3-Support): Providing expert-level troubleshooting to resolve complex technical issues, performing root-cause analysis and proposing solutions to prevent recurring failures.
  • Performance Optimisation & Load Testing: Optimizing system performance and scalability by assessing responsiveness under load and fine-tuning configurations to ensure a seamless experience for our users.
  • Security Management & Vulnerability Scanning: Proactively identifying and mitigating security vulnerabilities within the application and its environment, following a standardized framework.
  • Compliance & Audit Readiness (ISO 27001, SOC2, GDPR): Ensuring all operations meet legal and industry standards and providing the necessary evidence for internal or external audits.

What they're looking for

  • Proficiency in implementing and managing dashboards and alerting frameworks (e.g., Prometheus, Grafana, or New Relic).
  • Experience in Cloud-native environments (AWS), working with Infrastructure as Code and maintaining CI/CD pipelines.
  • Familiarity with Business Continuity, Disaster Recovery, and load / performance testing.
  • A curious mindset, eager to investigate root causes and find sustainable solutions.
  • A collaborative approach; you enjoy bridging the gap between technical teams and stakeholders to foster operational excellence.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

THE ROLE AND THE TEAM

The team serves as the bridge between Customer Operations, Software Development and Infrastructure teams and are responsible for ensuring stability, high-performance, and scalability across our global application landscape. 

Central to the role is proactive prevention: monitoring and alerting to catch issues before impact, proactive issue detection and rigorous Root Cause Analysis (RCA). You will support our 1st and 2nd level support teams with deep-dive code/environment level troubleshooting and automation of repetitive tasks. 

INCLUSIVE BY DESIGN

If you think you have what it takes, we encourage you to apply even if you don't meet every single requirement. You may just be the right candidate for this or other roles!

At Zalando, our vision is to be the leading European technology platform for fashion and lifestyle – one that thrives on diversity and is truly inclusive by design. We believe that diverse teams fuel innovation and creativity, and we actively seek out talent from all backgrounds.

We actively seek to reduce bias in our hiring and employment processes, focusing on your qualifications, skills, and contributions. To support this, we kindly ask that you refrain from including personal details such as your photo, age, or marital status in your CV, ensuring a fair and equitable evaluation based solely on your abilities and potential.

We are committed to providing an exceptional and accessible candidate experience for everyone. If you require any accommodations to support you throughout the hiring process, please let us know – we are here to assist you.

Discover more about our commitment to creating a diverse and inclusive workplace: https://jobs.zalando.com/en/our-culture/diversity-and-inclusion

WHAT WE’D LOVE YOU TO DO (AND LOVE DOING) 

  • Monitoring, Alerting & Observability: Implementing tools and dashboards to track system health and business processes in real-time, ensuring the team is proactively notified of any technical or functional anomalies.

  • Incident & Problem Management (L3-Support): Providing expert-level troubleshooting to resolve complex technical issues, performing root-cause analysis and proposing solutions to prevent recurring failures.

  • Performance Optimisation & Load Testing: Optimizing system performance and scalability by assessing responsiveness under load and fine-tuning configurations to ensure a seamless experience for our users.

  • Security Management & Vulnerability Scanning: Proactively identifying and mitigating security vulnerabilities within the application and its environment, following a standardized framework.

  • Compliance & Audit Readiness (ISO 27001, SOC2, GDPR): Ensuring all operations meet legal and industry standards and providing the necessary evidence for internal or external audits.

WE’D LOVE TO MEET YOU IF 

  • Proficiency in implementing and managing dashboards and alerting frameworks (e.g., Prometheus, Grafana, or New Relic).

  • Experience in Cloud-native environments (AWS), working with Infrastructure as Code and maintaining CI/CD pipelines.

  • Familiarity with Business Continuity, Disaster Recovery, and load / performance testing.

  • A curious mindset, eager to investigate root causes and find sustainable solutions.

  • A collaborative approach; you enjoy bridging the gap between technical teams and stakeholders to foster operational excellence.

 

OUR OFFER

Tradebyte provides a range of benefits, here’s an overview of what you can expect. Ask your Talent Acquisition Partner to learn more about what we offer.

  • Employee shares program

  • 40% off fashion and beauty products sold and shipped by Zalando, 30% off Lounge by Zalando, discounts from external partners

  • 2 paid volunteering days a year

  • Hybrid working model - work where it works for you within Germany or the UK, with occasional office attendance required for moments that matter. 

  • Work from abroad for up to 30 working days a year

  • 27 days of vacation a year to start for full-time employees

  • Relocation assistance available (subject to prior agreement)

  • Family services, including counseling and support

  • Health and wellbeing options (including Wellhub, formerly Gympass)

  • Mental health support and coaching available

  • Drive your development with our training offerings and biannual peer-to-peer review

ABOUT TRADEBYTE

At Tradebyte, we work with the biggest players in e-commerce - from trendsetting fashion brands to major online retailers. Our goal is to create a workplace where everyone feels valued, supported, and empowered.

We offer flexible work schedules, professional development opportunities, and a real commitment to work-life balance. We embrace diverse perspectives and ensure that every voice is heard and respected, because you’re unique and that matters to us.

As we continue to grow, we’re looking for new colleagues who share our passion. Love what you do - do what you love. Join Tradebyte, an independent company within the Zalando Group!

E-commerce

Company

ZalandoE-commerce
Ansbach, Germany

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Zalando's careers site·first seen 7 Aug 2026·last verified 15 Sept 2026·How we source jobs

Similar jobs

  • Site Reliability Engineer -Cloudstore (m/f/d) at gridscaleKöln, Germany–match not yet calculated
  • Senior Platform & Reliability Engineer (all genders) at contaboRemote (Germany)–match not yet calculated
  • Site Reliability Engineer at trivagoDüsseldorf, Germany–match not yet calculated
  • System Reliability Engineer at sereactGermany–match not yet calculated
  • Site Reliability Engineer -Openstack (m/w/d) at gridscaleKöln, Germany–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in Germany
  • Systems Engineer jobs in Germany
  • DevOps Engineer jobs in Germany
  • Cloud Engineer jobs in Germany
  • Site Reliability Engineer jobs in United States
  • Site Reliability Engineer jobs in India