NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in India
2 months ago
Apply with autofill
Apply with autofill
Skit.ai·SaaS·2 months ago
2 months ago

Site Reliability Engineer — Multi-Cloud Infrastructure

Bengaluru, IndiaFull-timeMid · 2-5 yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at Skit.ai

How you compare FREE

?
Your scoreYour score: not yet known
→
76
Top 10%Top 10%: 76 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in India.

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

About this role

About the Role
Skit.ai is the pioneer Conversational AI company transforming collections with omnichannel GenAI-powered assistants. Skit.ai’s Collection Orchestration Platform, the world’s first solution, streamlines collection conversations by syncing channels and accounts. Skit.ai’s Large Collection Model (LCM), a collection LLM, powers the strategy engine to optimize interactions, enhance customer experiences, and boost bottom lines for enterprises. Skit.ai has received several awards and recognitions, including the BIG AI Excellence Award 2024, Stevie Gold Winner 2023 for Most Innovative Company by The International Business Awards, and Disruptive Technology of the Year 2022 by CCW. Skit.ai is headquartered in New York City, NY. Visit https://skit.ai/

Job Title: Site Reliability Engineer — Multi-Cloud Infrastructure
Type: Full-time
Location: Bangalore

Why this role exists:
We run a voice AI platform for regulated enterprises in banking, telecom, and collections, spread across AWS, GCP, and Azure — for resilience, for cost, and because client data-residency rules leave us no choice. That's a lot of surface area: compute, networking, storage, identity, clusters, pipelines, and supporting services, all needing to stay healthy across three providers.
This role owns the day-to-day reliability and operations of that estate. It's the generalist counterpart to our real-time-platform SRE: where they go deep on the latency-critical call path, you go broad — keeping the whole infrastructure dependable, well-automated, and cost-sane, and sharing the on-call load. If you like knowing how everything fits together and making the boring parts reliable and self-serve, this is a good seat.

What you'll own:
  • Multi-cloud operations. Provision, operate, and keep healthy compute, networking, storage, and identity across AWS, GCP, and Azure — with sensible consistency instead of three snowflakes.
  • Infrastructure as code. Manage the estate through Terraform (or equivalent) and version control — reproducible environments, reviewed changes, no undocumented hand-tweaks.
  • CI/CD and delivery. Keep build and deploy pipelines fast and reliable so engineers ship safely and often.
  • Clusters and workloads. Run Kubernetes/container platforms and the supporting services (databases, queues, caches, internal tooling) that everything depends on.
  • Monitoring and on-call. Maintain monitoring and alerting for infrastructure health, take a turn in the rotation, and respond to and mitigate incidents with clear communication and blameless follow-up.
  • Cost and hygiene. Keep an eye on cloud spend, rightsizing, and waste; own the unglamorous but essential hygiene — patching, backups, secrets, and access.
  • Automation and toil reduction. Replace manual, repetitive operations with automation and self-service so the team scales without headcount scaling with it.

What the first year looks like:
  • First 90 days. Learn the estate across all three clouds. Take a turn on call. Close the most obvious gaps in monitoring, backups, and access hygiene.
  • By 6 months. More of the estate under consistent infrastructure-as-code. Reliable, reviewed CI/CD. A clearer, quieter alerting setup and documented runbooks for the common incidents.
  • By 12 months. Measurably less manual toil through automation and self-service. Sensible cost controls in place. Provisioning and environment setup that's repeatable rather than tribal knowledge.

What we're looking for:
Must-have
  • A few years in SRE, DevOps, or infrastructure operations for production systems, including on-call.
  • Hands-on experience across at least two of AWS, GCP, and Azure (all three is a strong plus).
  • Kubernetes and containers in production.
  • Infrastructure-as-code (Terraform or similar) and CI/CD pipelines.
  • Monitoring and alerting practice (e.g. Prometheus/Grafana) and structured incident handling.
  • A scripting/programming language for automation (Python, Go, or Bash beyond one-liners).
  • Solid Linux systems and networking fundamentals.
Nice-to-have
  • All three clouds run in production, and comfort designing for consistency across them.
  • Cost optimization / FinOps.
  • Secrets management, security hardening, and compliance/data-residency contexts.
  • PostgreSQL and other stateful-service operations at scale.
  • Some exposure to real-time or voice infrastructure — enough to back up the platform SRE on call.

Our Stack:
Representative — you'll help shape it. Multi-cloud across AWS, GCP, and Azure; Kubernetes/containers; Terraform and GitHub Actions CI/CD; PostgreSQL; Grafana/Tempo for monitoring; Modal for ML deployment; LiveKit/SIP telephony on the platform side.

How you'll know you're succeeding:
The infrastructure just works, across all three clouds, and when it doesn't it's caught early and fixed cleanly. Engineers provision what they need without filing tickets. Cloud spend is understood, not surprising. And the on-call rotation trends calmer because the estate is increasingly automated and self-healing.

We're an equal-opportunity employer and evaluate every candidate on merit. [Add benefits, compensation band, and application instructions before posting.]


SaaS

Company

Skit.aiSaaS
Bengaluru, India

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Skit.ai's careers site·first seen 5 Jul 2026·last verified 8 Sept 2026·How we source jobs

Similar jobs

  • AI/ML Site Reliability Engineer at London Stock Exchange Group (LSEG)Hyderabad, India–match not yet calculated
  • Site Reliability Engineer III at JPMorgan ChaseMumbai, India–match not yet calculated
  • Senior Site Reliability Engineer, Production Engineering at NVIDIABengaluru, India–match not yet calculated
  • Senior Mainframe Systems Programmer - Site Reliability Engineering at ensonoPune, India–match not yet calculated
  • Senior Site Reliability Engineer at Weekday AIBengaluru, India–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in India
  • DevOps Engineer jobs in India
  • Cloud Engineer jobs in India
  • Platform Engineer jobs in India
  • Site Reliability Engineer jobs in United States
  • Site Reliability Engineer jobs in United Kingdom