SRE II
Sign up free to see how well your resume matches this role.
What you'll do
- Build and maintain our in-house developer platforms — full-stack applications that provision and manage resources across multiple environments, spanning Kubernetes, CI/CD and observability.
- Run a multi-cluster Kubernetes fleet across multiple regions — several hundred production services, with complete GitOps delivery.
- Develop and implement automation for infrastructure provisioning and related tooling to streamline engineering and infrastructure operations.
- Design, implement and maintain infrastructure components.
- Identify, diagnose and resolve performance bottlenecks and issues across environments.
- Advocate for and promote best practices in software development and enhance the developer experience across engineering teams.
- Own scaling, capacity and application performance strategy — defining SLOs and SLIs, managing error budgets, and driving capacity planning across the fleet.
- Manage security and compliance policies as code across the fleet.
What they're looking for
- Bachelor's degree in computer science or a related field, or equivalent work experience.
- Proficiency in at least one modern programming language — Python preferred, for building full-scale application services covering web frameworks, databases and queues (not just scripting).
- In-depth experience with AWS, Linux and Kubernetes.
- Proven troubleshooting skills with the ability to diagnose and resolve complex issues.
- Experience with CI/CD and automation pipelines, and familiarity with infrastructure-as-code (Ansible, Terraform) and GitOps (ArgoCD) tooling.
- Sound reasoning about concurrency and failure: idempotency, optimistic concurrency, delivery semantics.
Nice to have
- Understanding of extending the Kubernetes API: CRDs and controller runtimes.
- Service mesh experience, ideally Istio.
- Familiarity with LLMs and MCP, fluency with AI coding tools, and experience authoring skills for coding agents.
- A strong point of view on where AI should and shouldn't be trusted in a safety-critical system
Summarised by NextRaise from the employer’s description, which follows in full below.
Full description from employer
Job Title: SRE II
Company: Upstox
Location: Mumbai
Work arrangement: 5 days in the office
About Upstox
At Upstox, we’re building the future of investing — simple, powerful, and for everyone. We're one of India’s fastest growing fintech platforms, backed by the best in the business, including Mr. Ratan Tata and Tiger Global, and on a mission to make wealth creation accessible to every Indian. From first-time investors to seasoned traders, millions trust us to power their financial journeys. We're not just moving fast — we’re moving with purpose. If you thrive in a high-energy, high-impact environment, you're in the right place.
What you'd own
- Build and maintain our in-house developer platforms — full-stack applications that provision and manage resources across multiple environments, spanning Kubernetes, CI/CD and observability.
- Run a multi-cluster Kubernetes fleet across multiple regions — several hundred production services, with complete GitOps delivery.
- Develop and implement automation for infrastructure provisioning and related tooling to streamline engineering and infrastructure operations.
- Design, implement and maintain infrastructure components.
- Identify, diagnose and resolve performance bottlenecks and issues across environments.
- Advocate for and promote best practices in software development and enhance the developer experience across engineering teams.
- Own scaling, capacity and application performance strategy — defining SLOs and SLIs, managing error budgets, and driving capacity planning across the fleet.
- Manage security and compliance policies as code across the fleet.
Required
- Bachelor's degree in computer science or a related field, or equivalent work experience.
- Proficiency in at least one modern programming language — Python preferred, for building full-scale application services covering web frameworks, databases and queues (not just scripting).
- In-depth experience with AWS, Linux and Kubernetes.
- Proven troubleshooting skills with the ability to diagnose and resolve complex issues.
- Experience with CI/CD and automation pipelines, and familiarity with infrastructure-as-code (Ansible, Terraform) and GitOps (ArgoCD) tooling.
- Sound reasoning about concurrency and failure: idempotency, optimistic concurrency, delivery semantics.
Nice to have
- Understanding of extending the Kubernetes API: CRDs and controller runtimes.
- Service mesh experience, ideally Istio.
- Familiarity with LLMs and MCP, fluency with AI coding tools, and experience authoring skills for coding agents.
- A strong point of view on where AI should and shouldn't be trusted in a safety-critical system
By applying for this position, you acknowledge that you have reviewed our Prospective Employee Privacy Notice, which outlines how Upstox collects, uses, and protects your Personal Information ("PI"). I accept Upstox's Prospective Employee Privacy Notice.
Upstox is an Equal Opportunity Employer; all qualified applicants will receive consideration for employment without regard to race, color, religion, gender, gender identity or expression, sexual orientation, national origin, genetics, disability, age, veteran status, or other characteristics.
Company
Company facts come from this company's own listings. We only show what the postings themselves carry.