Back to board
Early applicant
udchalo·1 day ago

Lead DevOps Engineer

Full-timeMid · 5+ yearsH1B likely

About this role

Lead DevOps Engineer
Job Description
Location: Pune
Experience: 5 years
Work Model: Work from office

About the role
We are looking for a hands-on Lead DevOps Engineer to lead the evolution of our cloud infrastructure and platform-engineering practices. This role combines technical leadership with direct execution: you will define platform decisions, turn them into dependable solutions, and own the infrastructure throughout its lifecycle.
Our platform is hosted on AWS. We are transitioning continuous integration and delivery workflows from Jenkins, AWS CodeBuild and AWS CodePipeline to GitHub Actions, while evolving our observability platform toward Dash0. The ideal candidate is an automation-first problem solver who can lead, build and troubleshoot with equal confidence.
Primary skills
•  AWS cloud infrastructure•  Platform engineering
•  Infrastructure as Code (Terraform)•  CI/CD and GitHub Actions
•  Infrastructure automation•  Observability (Dash0 / OpenTelemetry)
•  Linux and cloud networking•  Production troubleshooting
•  Technical decision-making•  End-to-end infrastructure ownership
Key responsibilities
  • Lead platform-engineering initiatives and define the technical direction for our infrastructure.
  • Own the design, provisioning, operation, security and continuous improvement of AWS infrastructure.
  • Evaluate technical options, make well-reasoned platform decisions, document them and execute them.
  • Drive the migration of CI/CD workflows from Jenkins, AWS CodeBuild and AWS CodePipeline to GitHub Actions.
  • Build reusable, reliable and secure delivery workflows that improve developer productivity.
  • Develop and maintain Infrastructure as Code and reduce manual infrastructure operations.
  • Lead the adoption of Dash0 and strengthen practices around metrics, logs, traces, dashboards and alerting.
  • Establish platform standards for reliability, scalability, security, observability and maintainability.
  • Troubleshoot infrastructure, deployment, networking and production reliability issues.
  • Replace repetitive operational work and recurring fixes with durable automation.
  • Partner with development teams to improve deployment processes and the overall developer experience.
  • Improve platform documentation, operational procedures and recovery practices.
Required qualifications and skills
  • Around five years of hands-on experience in DevOps, cloud infrastructure, site reliability or platform engineering.
  • Strong experience designing, deploying and operating production workloads on AWS.
  • Strong understanding of AWS networking, IAM, security, monitoring and infrastructure architecture.
  • Hands-on experience with Infrastructure as Code, preferably Terraform or an equivalent tool.
  • Strong experience designing and maintaining CI/CD pipelines, including practical GitHub Actions experience.
  • Experience with Jenkins, AWS CodeBuild or AWS CodePipeline and an understanding of migration considerations.
  • Good understanding of Linux, networking, DNS, TLS, access management and production troubleshooting.
  • Proficiency in infrastructure automation and scripting using Python, Bash or a similar language.
  • Understanding of observability across metrics, logs, distributed traces, dashboards and actionable alerting.
  • Ability to balance delivery speed, maintainability, reliability, security and cost when making decisions.
  • Strong written and verbal communication skills, with the ability to explain technical decisions clearly.


Preferred qualifications
  • Experience with Dash0 or another OpenTelemetry-based observability platform.
  • Experience migrating CI/CD or observability platforms.
  • Experience with Docker, Kubernetes, Helm, GitOps or Argo CD.
  • Familiarity with AWS cost management, governance, security controls and multi-account architecture.
  • Experience creating reusable platform components, deployment templates or standardized engineering workflows.
  • Experience mentoring engineers or leading cross-functional technical initiatives.
What we value
End-to-end ownershipTake responsibility for infrastructure throughout its lifecycle and stay involved when problems arise.Automation mindsetTurn repetitive, error-prone work into dependable tooling and workflows.
Hands-on leadershipSet technical direction while remaining comfortable implementing, testing and operating the solution.Structured troubleshootingInvestigate incidents methodically, identify root causes and implement lasting improvements.
Pragmatic decisionsEvaluate trade-offs and choose solutions that fit current needs and future growth.Continuous improvementProactively improve reliability, security, observability, delivery speed and developer experience.
What success looks like
  • Clear ownership and standards are established for the AWS platform.
  • Meaningful progress is made in migrating delivery workflows to GitHub Actions.
  • Infrastructure automation improves and reliance on manual operations declines.
  • Dash0 adoption advances and production visibility becomes stronger and more actionable.
  • Recurring operational issues are reduced through root-cause fixes and automation.
  • Deployment reliability and the infrastructure experience for development teams improve.
  • Technical documentation and repeatable operational practices are in place.