NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / DevOps Engineer in India
13 days ago
Apply with autofill
Apply with autofill
Skit.ai·SaaS·13 days ago
13 days ago

Lead DevOps Engineer

Bengaluru, IndiaFull-timeOn-siteSenior · 6+ yearsDevOps Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at Skit.ai

How you compare FREE

?
Your scoreYour score: not yet known
→
90
Top 10%Top 10%: 90 out of 100

Top 10% of NextRaise users matched against DevOps Engineer roles in India.

Must-have skills for this role

  • aws
  • gcp
  • azure
  • terraform

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Multi-cloud substrate: Production infrastructure across AWS, GCP, and Azure. Private interconnects (Direct Connect, Cloud Interconnect, ExpressRoute), transit/hub-spoke topologies, and cross-cloud latency managed as an explicit budget — p95 per hop in tens of milliseconds, not "best effort."
  • Real-time media plane: Self-hosted LiveKit and SIP infrastructure at scale. Media servers are stateful; you'll design session-affine, event-driven autoscaling (KEDA-class) that survives traffic tripling within an hour.
  • Model-serving infrastructure: GPU fleets (A100/H100/B200-class) for self-hosted ASR and open-weight LLMs — inference optimization, prefix caching, sticky-session routing, sub-500ms TTFT budgets — alongside managed APIs (Vertex AI/Gemini, Bedrock, Azure). Vendor failover is your design, not your incident.
  • Reliability & observability: OTel-native tracing (Grafana/Tempo stack), per-turn latency attribution across telephony/ASR/LLM/TTS, automated incident response and self-healing. You'll act as incident commander for infrastructure and write the runbooks you'd want at 3 a.m.
  • Cost engineering: Cost-per-minute is an SLO here. We cut per-minute serving cost ~18x in six months through caching, rightsizing, autoscaling, and workload re-architecture — you'll own the next 10x.
  • Security & compliance: Zero Trust across clouds: private endpoints/PrivateLink, IAM/RBAC, secrets management with rotation, WAF/DDoS protection. Operate controls for SOC 2 and ISO/IEC 27001; working command of ISO/IEC 42001:2023 (AI management systems) — hands-on preferred, rigorous theoretical grounding acceptable. You'll face bank and telecom auditors directly, including data-residency requirements.
  • Technical leadership: Terraform-first IaC standards, production-readiness reviews, mentoring SREs. "Lead" means you raise the floor of the whole team.

What they're looking for

  • 6+ years hands-on cloud infrastructure; 3+ years operating multiple clouds simultaneously in production; deep expertise in at least two of AWS/GCP/Azure
  • Real-time audio/video systems in production — WebRTC, SIP/PSTN, or streaming media; you've debugged jitter, not just read about it
  • Networking depth: VPC/VNet design, load balancing, DNS, NAT; private connectivity (Direct Connect / Cloud Interconnect / ExpressRoute, PrivateLink / Private Service Connect); transit gateways and cross-cloud mesh
  • Kubernetes at scale (EKS/GKE/AKS), Helm, and scaling stateful workloads; service mesh familiarity (Istio/Linkerd)
  • Infrastructure as Code: Terraform (non-negotiable) across multi-account/multi-project estates; drift is a bug
  • AI/ML serving in production: GPU allocation and scheduling, inference servers or serverless GPU platforms (vLLM / Triton / Modal / Baseten-class), streaming protocols (WebRTC, WebSocket, gRPC)
  • Production STT/TTS/LLM API operations: streaming integrations, quota management, multi-vendor failover (Deepgram / Google / Azure / Whisper-class ASR; ElevenLabs / Azure-class TTS)
  • Security fundamentals: IAM/RBAC, secrets management (Vault or cloud-native), encryption and key rotation
  • CI/CD: GitHub Actions or GitLab CI with security scanning integrated into the pipeline

Nice to have

  • LiveKit, pipecat, Twilio, or comparable real-time platforms; SIP trunking and PSTN integration
  • KEDA or other event-driven autoscaling used in anger
  • MLOps: model versioning, canary and A/B rollout
  • FinOps discipline: reserved/spot strategy, unit-economics reporting
  • Certifications: AWS SA Professional, GCP Professional Cloud Architect, Azure Solutions Architect Expert
  • ISO/IEC 42001:2023 exposure

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Job Title: Lead DevOps Engineer
Location: Bengaluru (100% WFO)
Job Type: Full-time

The problem
Skit.ai runs autonomous voice agents for regulated enterprises — India's largest banks and telcos, and US collections operations. Every call is a live distributed system: PSTN/SIP → media server → ASR → LLM → TTS → back, spread across three clouds and multiple vendors, with a conversational response budget measured in hundreds of milliseconds.
The platform peaks at roughly **1 million calls per hour**. Billed minutes grew **5,000x+ in eight months**. At this scale, infrastructure is not a support function — latency, cost-per-minute, and auditability are product features. When infra degrades, a customer mid-sentence hears silence.

We're hiring a Lead DevOps Engineer to own this substrate and keep it ahead of the growth curve.

What you'll own:
  • Multi-cloud substrate:  Production infrastructure across AWS, GCP, and Azure. Private interconnects (Direct Connect, Cloud Interconnect, ExpressRoute), transit/hub-spoke topologies, and cross-cloud latency managed as an explicit budget — p95 per hop in tens of milliseconds, not "best effort."
  • Real-time media plane: Self-hosted LiveKit and SIP infrastructure at scale. Media servers are stateful; you'll design session-affine, event-driven autoscaling (KEDA-class) that survives traffic tripling within an hour.
  • Model-serving infrastructure: GPU fleets (A100/H100/B200-class) for self-hosted ASR and open-weight LLMs — inference optimization, prefix caching, sticky-session routing, sub-500ms TTFT budgets — alongside managed APIs (Vertex AI/Gemini, Bedrock, Azure). Vendor failover is your design, not your incident.
  • Reliability & observability: OTel-native tracing (Grafana/Tempo stack), per-turn latency attribution across telephony/ASR/LLM/TTS, automated incident response and self-healing. You'll act as incident commander for infrastructure and write the runbooks you'd want at 3 a.m.
  • Cost engineering: Cost-per-minute is an SLO here. We cut per-minute serving cost ~18x in six months through caching, rightsizing, autoscaling, and workload re-architecture — you'll own the next 10x.
  • Security & compliance: Zero Trust across clouds: private endpoints/PrivateLink, IAM/RBAC, secrets management with rotation, WAF/DDoS protection. Operate controls for SOC 2 and ISO/IEC 27001; working command of ISO/IEC 42001:2023 (AI management systems) — hands-on preferred, rigorous theoretical grounding acceptable. You'll face bank and telecom auditors directly, including data-residency requirements.
  • Technical leadership: Terraform-first IaC standards, production-readiness reviews, mentoring SREs. "Lead" means you raise the floor of the whole team.

Problems on our plate right now
  • Scaling stateful, self-hosted media servers past current concurrency ceilings — HPA on CPU doesn't cut it
  • Migrating LLM inference from managed APIs to self-hosted open-weight models on GPUs without breaking TTFT budgets
  • ASR, LLM, and telephony living in different clouds: interconnect topology that keeps the packet path short and private
  • Multi-region DR that satisfies bank audits without doubling spend
If these read as interesting rather than terrifying, keep reading.
Must-have
  • 6+ years  hands-on cloud infrastructure; 3+ years operating multiple clouds simultaneously in production; deep expertise in at least two of AWS/GCP/Azure
  • Real-time audio/video systems in production — WebRTC, SIP/PSTN, or streaming media; you've debugged jitter, not just read about it
  • Networking depth: VPC/VNet design, load balancing, DNS, NAT; private connectivity (Direct Connect / Cloud Interconnect / ExpressRoute, PrivateLink / Private Service Connect); transit gateways and cross-cloud mesh
  • Kubernetes at scale (EKS/GKE/AKS), Helm, and scaling *stateful* workloads; service mesh familiarity (Istio/Linkerd)
  • Infrastructure as Code: Terraform (non-negotiable) across multi-account/multi-project estates; drift is a bug
  • AI/ML serving in production: GPU allocation and scheduling, inference servers or serverless GPU platforms (vLLM / Triton / Modal / Baseten-class), streaming protocols (WebRTC, WebSocket, gRPC)
  • Production STT/TTS/LLM API operations: streaming integrations, quota management, multi-vendor failover (Deepgram / Google / Azure / Whisper-class ASR; ElevenLabs / Azure-class TTS)
  • Security fundamentals: IAM/RBAC, secrets management (Vault or cloud-native), encryption and key rotation
  • CI/CD: GitHub Actions or GitLab CI with security scanning integrated into the pipeline

Strong signal (nice-to-have)
  • LiveKit, pipecat, Twilio, or comparable real-time platforms; SIP trunking and PSTN integration
  • KEDA or other event-driven autoscaling used in anger
  • MLOps: model versioning, canary and A/B rollout
  • FinOps discipline: reserved/spot strategy, unit-economics reporting
  • Certifications: AWS SA Professional, GCP Professional Cloud Architect, Azure Solutions Architect Expert
  • ISO/IEC 42001:2023 exposure
What we're NOT looking for
  • Single-cloud depth with documentation-level knowledge of the other two
  • Tool-checklist DevOps without production AI/ML serving scars
  • "Can learn quickly" as the primary qualification — this role needs day-one production credibility
  • Anyone who has never traced a packet across a cloud boundary

SaaS

Company

Skit.aiSaaS
Bengaluru, India

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Skit.ai's careers site·first seen 7 Sept 2026·last verified 8 Sept 2026·How we source jobs

Similar jobs

  • Senior DevOps Engineer at webMumbai, India–match not yet calculated
  • AWS Devops at synechronHyderabad, India–match not yet calculated
  • SR DevOps Engineer at allegionBengaluru, India–match not yet calculated
  • Senior R&D Engineer DevOps - IDC at HitachiBengaluru, India–match not yet calculated
  • DevOps at tradelabBengaluru, India–match not yet calculated

Browse more jobs

  • DevOps Engineer jobs in India
  • Cloud Engineer jobs in India
  • Platform Engineer jobs in India
  • Systems Engineer jobs in India
  • DevOps Engineer jobs in United States
  • DevOps Engineer jobs in France