Site Reliability Engineer
Sign up free to see how well your resume matches this role.
What you'll do
- Lead incident response, conduct RCAs, and design full-stack observability (Prometheus, Grafana, Datadog) to eliminate alert fatigue.
- Build AWS cloud-native infrastructure using Terraform and integrate security tools (SAST/DAST) into CI/CD pipelines.
- Mentor junior SREs, contribute to on-call rotations, and represent SRE in architecture reviews.
- Manage Cloudflare CDN and Workers.
What they're looking for
- 6 to 10 Years of experience
- Must leverage AIOps platforms for anomaly detection and utilize LLM-powered tools to accelerate MTTR
- Must own end-to-end reliability of AEM environments (Author, Publish, Dispatcher, and AEMaaCS) and manage OSGi configurations and replication queues
Summarised by NextRaise from the employer’s description, which follows in full below.
Full description from employer
Experience required - 6 to 10 Years
- Reliability & Observability: Lead incident response, conduct RCAs, and design full-stack observability (Prometheus, Grafana, Datadog) to eliminate alert fatigue.
- AI Adoption (Mandatory): Must leverage AIOps platforms for anomaly detection and utilize LLM-powered tools to accelerate MTTR.
- AEM Management (Mandatory): Must own end-to-end reliability of AEM environments (Author, Publish, Dispatcher, and AEMaaCS) and manage OSGi configurations and replication queues.
- Infrastructure & Security: Build AWS cloud-native infrastructure using Terraform and integrate security tools (SAST/DAST) into CI/CD pipelines.
- Leadership: Mentor junior SREs, contribute to on-call rotations, and represent SRE in architecture reviews.
- CDN: Manage Cloudflare CDN and Workers.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.