Senior Devops Engineer
Mumbai, IndiaFull-timeMid · 5-8 years
About this role
- Infrastructure Management (AWS)
- Own and managethe entire AWS infrastructure — EC2, RDS, S3, ECS/EKS,VPC, IAM, CloudFront, Route 53, and related services.
- Design for high availability and fault tolerance; ensure infra can handle payment-grade uptime requirements.
- Right-size and optimise infrastructure for cost without compromising reliability.
- Maintain environment parity across dev, staging, and production.
- CI/CD Pipelines
- Build, maintain, and continuously improveCI/CD pipelines acrossall services supporting a daily release cadence.
- Manage deployment triggering across multiple AWS availability zones;handle traffic routingbetween zones including manual intervention when required.
- Ensure fast, reliable, and safe deployments with rollback capabilities and pre/post-release health checks.
- Work closely with the engineering team to reduce deployment friction and maintain release velocity.
- Manage branching strategies, environment promotion, and deployment gates.
- Monitoring, Alerting & Incident Response
- Own site monitoring end-to-end — set up and maintain CloudWatch and Site24x7 (or equivalent) for uptime checks, dashboards, log aggregation, and alerting.
- Configure intelligent alertingrules that surfacereal issues withoutcreating noise; maintainrunbooks for common alert scenarios.
- Be the firstresponder for production alerts — acknowledge, triage, and resolveor escalate withindefined SLAs. Proactive action on alerts is a core expectation of this role.
- Conduct root cause analysis (RCA) for all production incidents and drive fixes to prevent recurrence.
- Maintain an on-call schedule; this role requires availability outside business hours for critical alerts.
- System Automation
- Automate repetitive operational tasks — provisioning, scaling, patching, backups,and configuration management.
- Manage infrastructure-as-code using Terraform or equivalent; no manual console changes in production.
- Automate monitoring setup, alerting thresholds, and runbook execution wherever possible.
- Security & Compliance
- Implement and maintainsecurity rules, firewallpolicies, network ACLs,and IAM roleswith least-privilege principles.
- Manage VPN setup and access controls for internal teams and vendors.
- Manage access provisioning and deprovisioning for all team members across AWS and related services.
- Ensure infrastructure compliance with PCI DSS and ISO 27001 requirements — work closelywith the InfoSec Lead on audit evidence and remediation.
- Manage SSL/TLS certificates, key rotation, secretsmanagement (AWS SecretsManager / Vault),and encryption at rest and in transit.
- Releases & Server Operations
- Coordinate and executedaily production releases— validate pre-release checklists, trigger deployments across zones, manage traffic routing, and monitor post-release health.
- Handle server maintenance, OS upgrades, dependency patching, and scheduled downtime windows.
- Manage database backups, restoration drills, and disaster recovery procedures.
- Maintain up-to-date infrastructure documentation and runbooks.
- AI-Augmented Development Infrastructure
- Design and manage infrastructure for AI-assisted development workflows — including on-demand provisioningand teardown of ephemeral EC2 or container instances used by AI dev tools such as Claude Code.
- Integrate AI dev tooling into CI/CD pipelines — enabling automated code generation, review,and testing stages that spin up isolated compute, execute tasks, and clean up on completion.
- Implement IAM policies, network boundaries, and cost guardrails for ephemeral AI development instances.
- Build monitoring and observability into AI-augmented pipelines— tracking instancelifecycle, run times, failure rates, and compute costs.
