NextRaise Logo
NextRaise
JobsDiscover rolesJob TrackerTrack applied positionsMy ResumesBuild & optimize resumes
Tools
Job Match AnalyzerPaste JD, get fit scoreATS ScoreScan for ATS issues
Chrome Extension
ResumesJobsProfile
Jobs / Site Reliability Engineer in India
4 hours agoBe an early applicant
Skyflow·4 hours ago
4 hours agoBe an early applicant

Site Reliability Engineer/Cloud Platform Engineer - Operations (PST Timezone)

IndiaFull-timeRemoteMid · 4+ years

Sign up free to see how well your resume matches this role.

About this role

Skyflow secures the flow of data across datastores, models, and agents. Enterprises turn to Skyflow as their runtime AI data control layer to protect sensitive data, enable safe AI deployment, and unlock full value from their applications, data platforms, and AI systems.

Skyflow is trusted by Fortune 500 enterprises, and leading SaaS companies in financial services, healthcare, retail, travel and hospitality.

Skyflow is headquartered in Palo Alto, California and was founded in 2019. For more information, visit www.skyflow.com or follow on X and LinkedIn.

 

About the role

We're looking for a Platform Engineer to help build and run the infrastructure that powers a multi-tenant, security-sensitive B2B cloud product used by enterprise customers around the world. Our platform spans multi-tenant environments and single-tenant (BYOC) deployments into customer cloud accounts, so the bar for automation, consistency, and reliability is high — every environment needs to be provisioned, upgraded, monitored, and recovered the same way, at scale, with minimal manual intervention.

This is a hands-on engineering role, not a ticket-queue ops role. You'll spend most of your time writing Go and Python to turn repeatable infrastructure work into platform capabilities — provisioning pipelines, internal CLIs, self-service tooling, and automation that lets the rest of engineering ship without waiting on infrastructure as a bottleneck. You'll also carry deep operational ownership: on-call, incident response, and root-cause analysis on the systems you build.

 

You have:

  • 4+ years of experience in platform engineering, infrastructure engineering, DevOps, or SRE roles, with real ownership of production cloud infrastructure.

  • Strong software engineering skills in Go and/or Python — you write tested, maintainable code and think of infrastructure automation as software, not scripting.

  • Deep hands-on experience with Kubernetes in production (workload scheduling, networking, autoscaling, upgrades) and Infrastructure-as-Code tooling such as Terraform, Pulumi, or OpenTofu.

  • Solid experience with at least one major public cloud (AWS or GCP); experience operating in both is a strong plus given our multi-cloud footprint.

  • Track record of building tools or platforms that other engineers use — internal CLIs, provisioning frameworks, self-service portals, or automation pipelines — not just maintaining existing infrastructure.

  • Comfort operating in a B2B environment with enterprise customers, where infrastructure changes carry compliance, security, and contractual weight (e.g. dedicated/BYOC deployments, uptime SLAs, audit requirements).

  • Strong incident-response instincts: you can debug distributed systems under pressure, drive a root-cause analysis to a real fix, and communicate clearly during and after an incident.

  • A bias toward root-causing and automating away recurring problems over repeatedly firefighting the same issue.

  • Lead, mentor, and technically guide a team of junior/early-career platform engineers.

 

You will:

  • Provide operational support aligned with US time zones, ensuring system reliability and availability.

  • Design and build automation that provisions and manages cloud infrastructure end-to-end — new environment onboarding, upgrades, scaling, and decommissioning — across multiple cloud providers (AWS, GCP) and multiple deployment models (multi-tenant and dedicated/BYOC).

  • Write production-grade Go and Python services and CLIs that turn infrastructure operations into self-service platform capabilities for other engineering teams, rather than one-off scripts or manual runbooks.

  • Own and evolve the Infrastructure-as-Code stack (Terraform/OpenTofu, Helm, GitOps/ArgoCD) that defines every environment, and drive migrations across the fleet (Kubernetes version upgrades, node pool migrations, service mesh changes) with minimal customer impact.

  • Operate and scale core platform services — Kubernetes clusters, service mesh (Istio), data stores (Aerospike, PostgreSQL), messaging (Kafka), and GPU-backed inference workloads — with a focus on capacity planning, cost efficiency, and right-sizing.

  • Build and improve observability and alerting (metrics, logs, synthetic monitoring) so that failures are caught before customers notice, and drive the automation that turns repeat incidents into permanent fixes.

  • Participate in an on-call rotation, lead incident response for the systems you own, and write root-cause analyses that result in concrete corrective action — not just documentation.

  • Partner with security and compliance stakeholders to build guardrails (secrets management, access control, network policy, audit logging) directly into the platform, so secure-by-default is the path of least resistance for every team.

  • Continuously identify manual, repetitive, or error-prone infrastructure work and eliminate it — the measure of success in this role is less manual toil across the org, not more tickets closed.

 

Nice to have

  • Experience with service mesh (Istio/Envoy), GitOps workflows (ArgoCD/Flux), or policy-as-code (OPA).

  • Experience operating stateful systems in production — distributed databases (Aerospike, PostgreSQL, Cassandra), message queues (Kafka), or GPU-backed inference workloads.

  • Experience with security-sensitive or regulated environments (PCI, SOC 2, HIPAA) and building compliance controls into infrastructure by default.

  • Experience with cost optimization and capacity planning at fleet scale (right-sizing, autoscaling policy design, spot/preemptible usage).

  • Contributions to open-source infrastructure tooling, or experience building an internal developer platform (IDP).

 

Our stack (roughly)

Go, Python · Kubernetes, Helm, ArgoCD · Terraform/OpenTofu · AWS, GCP · Istio · Aerospike, PostgreSQL, Kafka · Coralogix/Prometheus-style observability · GitHub Actions and Spacelift-style CI/CD for infrastructure changes.

 

Benefits:

  • Work from home expense

  • Excellent Health Insurance Options

  • Very generous PTO

  • Flexible Hours

  • Generous Equity

At Skyflow, we believe that diverse teams are the strongest teams. We invite applicants of all genders, races, ethnicities, nationalities, ages, religions, sexual orientations, disability statuses, educational experiences, family situations, and socio-economic backgrounds.

 
 
AI tools
Apply faster with autofillThe NextRaise extension autofills your application in one click.Get the extension

Similar jobs

  • Senior Site Reliability Engineer/Cloud Platform Engineer at skyflowIndia
  • Principal Site Reliability Engineer at servicetitanBengaluru, India
  • Senior Site Reliability Engineer at rapidaiBengaluru, India
  • Director, Site Reliability Engineering at OktaBengaluru, India
  • AI/MLOps SRE Lead Engineer at regeneronHyderabad, India
  • Senior PostgreSQL SRE at BarclaysPune, India