Back to board
Early applicant
skanai·3 hours ago

Sr Site Reliability Engineer

Bengaluru, IndiaFull-timeMid · 5-8 years

About this role

Be at the Forefront of the Agentic AI Revolution
At Skan AI, you'll be part of the team pioneering the context engine for human and agentic execution, bringing context from enterprise operators, systems, and processes to power how the world's largest organizations execute their most complex, mission-critical work.
Why Join Skan AI
We're in hyper-growth mode at exactly the right moment in history. As enterprises race to adopt agentic AI, we're uniquely positioned to deliver the clear signal they desperately need: a platform that trains and grounds AI Agents in trillions of real execution signals, enabling reliable, compliant automation of their most complex processes.
Backed by Dell Technologies Capital and other leading investors, we're the only company that can bridge the gap between AI's promise and enterprise reality, making us perfectly positioned to define the agentic era for modern enterprises.
Our diverse, collaborative team of 250+ innovators is solving category-defining challenges at the intersection of AI, process intelligence, and enterprise work. Diverse perspectives fuel breakthrough thinking, cross-functional collaboration is the norm, and our work directly transforms how Fortune 500 companies operate. We are shaping the future of work itself.

The Site Reliability Engineer (SRE) owns the reliability, availability, and operational health of Skan's cloud-hosted customer environments. Sitting within the Customer Cloud Ops team, this role bridges software and infrastructure — applying engineering discipline to automate toil, respond to incidents rapidly, and uphold the SLA commitments that protect Skan's enterprise relationships.
The SRE is accountable for keeping production customer environments running at the performance and availability standards enterprise clients expect. This means being on-call, being proactive about operational risk, and continuously eliminating the manual work that gets in the way of reliable operations.
WHY THIS ROLE EXISTS
As Skan's cloud-hosted customer base grows, the complexity and volume of environment management, incident response, and reliability engineering grows with it. Without dedicated SRE capability, operational toil accumulates, incidents take longer to resolve, and SLA commitments are at risk — damaging customer trust and creating costly escalations.
The SRE function applies software engineering practices to infrastructure and operations — replacing manual, reactive processes with automation, runbooks, and agentic workflows that scale.
KEY RESPONSIBILITIES
  • Monitor platform health and performance across all cloud-hosted customer environments using observability tooling (Prometheus, Grafana, Datadog, or equivalent)
  • Respond to and own P1/P2 incidents — lead triage, diagnosis, and resolution; drive MTTR reduction through structured post-incident review
  • Perform ongoing capacity and environment planning for customer cloud deployments — anticipating growth and preventing resource-related incidents
  • Design and implement SRE automation to eliminate repetitive operational toil — agentic alerting, auto-remediation scripts, and automated runbook execution
  • Manage change and release events that affect production customer environments — coordinating with DevOps and product teams to minimize risk
  • Maintain and improve runbooks for all known failure patterns and operational procedures
  • Contribute to HA/DR playbook validation
  • Participate in on-call rotation and respond to alerts within defined SLA windows
  • Track and report on SLO/SLA adherence — contributing to monthly operational reports and QBR data
  • Identify and escalate environment risks proactively — before they become customer-facing incidents
  • Collaborate with the Automation Engineering team to develop agentic workflows that automate triage, routing, and remediation
KEY DELIVERABLES
  • Incident response records and post-mortems for all P1/P2 events — with root cause, remediation steps, and prevention actions documented
  • SLA/SLO dashboards: availability and performance reports for all cloud customer environments — updated continuously
  • Capacity and environment sizing plans per customer — reviewed quarterly or on significant usage change
  • Automated toil-reduction scripts and agentic remediation workflows — measurable reduction in manual operational hours
  • Runbooks for all known failure patterns — maintained and validated against real incidents
  • Change management records for all production environment events
  • Target: ≥ 99.9% uptime across all cloud-hosted customer environments
KEY SKILLS & QUALIFICATIONS
  • 5-8 years of experience as a SRE engineer
  • Cloud platforms: AWS, Azure, or GCP — environment management, networking, IAM, and observability at scale
  • Observability and monitoring: Prometheus, Grafana, Datadog, or equivalent — building dashboards, alerts, and SLO tracking
  • Infrastructure as Code: Terraform, Ansible, or Pulumi — provisioning and configuration management
  • Incident management: on-call discipline, structured MTTR mindset, post-mortem culture, and blameless review practices
  • Scripting and automation: Python, Bash — automation of operational tasks and agentic workflow development
  • Linux systems administration: process management, log analysis, performance tuning
  • Container orchestration: Kubernetes and Docker — deployment management and debugging in production
  • SRE fundamentals: SLI/SLO/SLA definition, error budget management, toil measurement and reduction
  • Strong written documentation skills — clear, evidence-based runbooks and incident reports

Skan AI is an equal opportunity employer committed to building a diverse, inclusive, and respectful workplace around the world. We do not discriminate based on race, color, religion or belief, sex (including pregnancy, sexual orientation, gender identity, or gender expression), national origin, ancestry, age, disability, medical condition, genetic information, marital or family status, military or veteran status, or any other characteristic protected by applicable laws in the locations where we operate.
We welcome people from all backgrounds and provide reasonable accommodations throughout the hiring process.