Back to board
relanto·1 month ago

Senior DevOps Engineer

Fremont, United States of AmericaFull-timeMid · 5+ yearsH1B likely

About this role

Responsibilities:
  • Manage, administer, and optimize production Kubernetes clusters across Amazon EKS and Google Anthos/GKE environments.
  • Perform cluster lifecycle management including upgrades, patching, capacity planning, autoscaling, and performance optimization.
  • Design and implement Kubernetes security best practices, including RBAC, network policies, pod security standards, and cluster hardening.
  • Configure and maintain ingress controllers such as NGINX Ingress and AWS ALB Ingress Controller for secure and scalable application delivery.
  • Build and maintain observability platforms using Prometheus, Grafana, Splunk, and cloud-native monitoring services for metrics, alerting, and log analytics.
  • Develop and manage GitOps-driven CI/CD pipelines using ArgoCD, Jenkins, GitHub Actions, and Bitbucket to support automated deployments and release management.
  • Provision and manage cloud infrastructure using Terraform, creating reusable modules and enforcing Infrastructure as Code best practices.
  • Collaborate with development, security, and platform engineering teams to streamline application deployments and operational processes.
  • Manage and optimize AWS infrastructure components including networking, compute, storage, security, and monitoring services.
  • Troubleshoot production incidents, perform root cause analysis, and implement preventive measures to improve platform reliability and availability.
  • Define and enforce DevOps standards, operational procedures, and cloud governance best practices across distributed engineering teams.

Required Skills:
  • 5+ years of hands-on experience managing Kubernetes clusters in production environments.
  • Strong expertise with Kubernetes administration, including EKS, GKE, or Anthos platforms.
  • Experience with Kubernetes networking, ingress controllers, service meshes, RBAC, autoscaling, and security hardening.
  • Hands-on experience building and managing CI/CD pipelines using tools such as ArgoCD, Jenkins, GitHub Actions, and Bitbucket.
  • Strong proficiency in Terraform and Infrastructure as Code (IaC) methodologies.
  • Deep understanding of AWS services including VPC, IAM, EC2, S3, Route 53, ELB/ALB, ECR, CloudWatch, KMS, and Secrets Manager.
  • Experience implementing monitoring, logging, and observability solutions using Prometheus, Grafana, Splunk, and cloud-native monitoring tools.
  • Strong understanding of containerization technologies including Docker and Kubernetes ecosystem components.
  • Experience with Linux system administration, shell scripting, and automation.
  • Strong troubleshooting, performance tuning, and incident management skills.
  • Excellent communication skills with the ability to collaborate effectively across engineering and business teams.
  • Experience working in Agile environments and supporting distributed delivery teams.