Back to board
getshopse·20 days ago

Senior Devops Engineer

Mumbai, IndiaFull-timeMid · 5-8 years

About this role

  1. Infrastructure Management (AWS)
    • Own and managethe entire AWS infrastructure — EC2, RDS, S3, ECS/EKS,VPC, IAM, CloudFront, Route 53, and related services.
    • Design for high availability and fault tolerance; ensure infra can handle payment-grade uptime requirements.
    • Right-size and optimise infrastructure for cost without compromising reliability.
    • Maintain environment parity across dev, staging, and production.

  1. CI/CD Pipelines
    • Build, maintain, and continuously improveCI/CD pipelines acrossall services supporting a daily release cadence.
    • Manage deployment triggering across multiple AWS availability zones;handle traffic routingbetween zones including manual intervention when required.
    • Ensure fast, reliable, and safe deployments with rollback capabilities and pre/post-release health checks.
    • Work closely with the engineering team to reduce deployment friction and maintain release velocity.
    • Manage branching strategies, environment promotion, and deployment gates.

  1. Monitoring, Alerting & Incident Response
    • Own site monitoring end-to-end — set up and maintain CloudWatch and Site24x7 (or equivalent) for uptime checks, dashboards, log aggregation, and alerting.

    • Configure intelligent alertingrules that surfacereal issues withoutcreating noise; maintainrunbooks for common alert scenarios.
    • Be the firstresponder for production alerts — acknowledge, triage, and resolveor escalate withindefined SLAs. Proactive action on alerts is a core expectation of this role.
    • Conduct root cause analysis (RCA) for all production incidents and drive fixes to prevent recurrence.
    • Maintain an on-call schedule; this role requires availability outside business hours for critical alerts.

  1. System Automation
    • Automate repetitive operational tasks — provisioning, scaling, patching, backups,and configuration management.
    • Manage infrastructure-as-code using Terraform or equivalent; no manual console changes in production.
    • Automate monitoring setup, alerting thresholds, and runbook execution wherever possible.

  1. Security & Compliance
    • Implement and maintainsecurity rules, firewallpolicies, network ACLs,and IAM roleswith least-privilege principles.
    • Manage VPN setup and access controls for internal teams and vendors.
    • Manage access provisioning and deprovisioning for all team members across AWS and related services.
    • Ensure infrastructure compliance with PCI DSS and ISO 27001 requirements — work closelywith the InfoSec Lead on audit evidence and remediation.
    • Manage SSL/TLS certificates, key rotation, secretsmanagement (AWS SecretsManager / Vault),and encryption at rest and in transit.
  2. Releases & Server Operations
    • Coordinate and executedaily production releases— validate pre-release checklists, trigger deployments across zones, manage traffic routing, and monitor post-release health.
    • Handle server maintenance, OS upgrades, dependency patching, and scheduled downtime windows.
    • Manage database backups, restoration drills, and disaster recovery procedures.
    • Maintain up-to-date infrastructure documentation and runbooks.

  1. AI-Augmented Development Infrastructure
    • Design and manage infrastructure for AI-assisted development workflows — including on-demand provisioningand teardown of ephemeral EC2 or container instances used by AI dev tools such as Claude Code.
    • Integrate AI dev tooling into CI/CD pipelines — enabling automated code generation, review,and testing stages that spin up isolated compute, execute tasks, and clean up on completion.
    • Implement IAM policies, network boundaries, and cost guardrails for ephemeral AI development instances.
    • Build monitoring and observability into AI-augmented pipelines— tracking instancelifecycle, run times, failure rates, and compute costs.