Senior DevOps Engineer
Fremont, United States of AmericaFull-timeMid · 5+ yearsH1B likely
About this role
Responsibilities:
- Manage, administer, and optimize production Kubernetes clusters across Amazon EKS and Google Anthos/GKE environments.
- Perform cluster lifecycle management including upgrades, patching, capacity planning, autoscaling, and performance optimization.
- Design and implement Kubernetes security best practices, including RBAC, network policies, pod security standards, and cluster hardening.
- Configure and maintain ingress controllers such as NGINX Ingress and AWS ALB Ingress Controller for secure and scalable application delivery.
- Build and maintain observability platforms using Prometheus, Grafana, Splunk, and cloud-native monitoring services for metrics, alerting, and log analytics.
- Develop and manage GitOps-driven CI/CD pipelines using ArgoCD, Jenkins, GitHub Actions, and Bitbucket to support automated deployments and release management.
- Provision and manage cloud infrastructure using Terraform, creating reusable modules and enforcing Infrastructure as Code best practices.
- Collaborate with development, security, and platform engineering teams to streamline application deployments and operational processes.
- Manage and optimize AWS infrastructure components including networking, compute, storage, security, and monitoring services.
- Troubleshoot production incidents, perform root cause analysis, and implement preventive measures to improve platform reliability and availability.
- Define and enforce DevOps standards, operational procedures, and cloud governance best practices across distributed engineering teams.
Required Skills:
- 5+ years of hands-on experience managing Kubernetes clusters in production environments.
- Strong expertise with Kubernetes administration, including EKS, GKE, or Anthos platforms.
- Experience with Kubernetes networking, ingress controllers, service meshes, RBAC, autoscaling, and security hardening.
- Hands-on experience building and managing CI/CD pipelines using tools such as ArgoCD, Jenkins, GitHub Actions, and Bitbucket.
- Strong proficiency in Terraform and Infrastructure as Code (IaC) methodologies.
- Deep understanding of AWS services including VPC, IAM, EC2, S3, Route 53, ELB/ALB, ECR, CloudWatch, KMS, and Secrets Manager.
- Experience implementing monitoring, logging, and observability solutions using Prometheus, Grafana, Splunk, and cloud-native monitoring tools.
- Strong understanding of containerization technologies including Docker and Kubernetes ecosystem components.
- Experience with Linux system administration, shell scripting, and automation.
- Strong troubleshooting, performance tuning, and incident management skills.
- Excellent communication skills with the ability to collaborate effectively across engineering and business teams.
- Experience working in Agile environments and supporting distributed delivery teams.
