Back to board
saarthee·3 months ago

Senior Site Reliability Engineer

Full-timeSenior · 6+ yearsH1B likely

About this role

About Saarthee:

Saarthee is a Global Strategy, Analytics, Technology and AI consulting company, where our passion for helping others fuels our approach and our products and solutions. We are a onestop shop for all things data and analytics. Unlike other analytics consulting firms that are technology or platform specific, Saarthee’s holistic and tool agnostic approach is unique in the marketplace. Our Consulting Value Chain framework meets our customers where they are in their data journey. Our diverse and global team work with one objective in mind: Our Customers’ Success. At Saarthee, we are passionate about guiding organizations towards insights fueled success. That’s why we call ourselves Saarthee–inspired by the Sanskrit word ‘Saarthi’, which means charioteer, trusted guide, or companion. Cofounded in 2015 by Mrinal Prasad and Shikha Miglani, Saarthee already encompasses all the components of Data Analytics consulting. Saarthee is based out of Philadelphia, USA with office in UK and India

Position Summary:

We are looking for a Senior Site Reliability Engineer (SRE) with deep expertise in observability, cloud-native infrastructure, and large-scale distributed systems. This role is highly hands-on and focuses on designing, building, and operating reliable, observable, and scalable platforms running on Kubernetes, with a strong preference for Google Cloud Platform (GCP) and AWS.

Your Role Responsibilities and Duties:

  • Design, implement, and operate highly available and resilient Kubernetes-based systems.
  • Define, monitor, and enforce SLIs, SLOs, and error budgets to ensure service reliability.
  • Lead incident response, root cause analysis (RCA), and postmortems, driving continuous improvement.
  • Architect and manage observability platforms for metrics, logging, tracing, and alerting.
  • Work hands-on with Prometheus, Alertmanager, OpenTelemetry, Grafana, and Loki / ELK / OpenSearch.
  • Implement cloud-native monitoring and logging, with preference for GCP Cloud Monitoring & Logging.
  • Establish actionable alerting standards to reduce noise and improve response effectiveness.
  • Build and manage cloud infrastructure on GCP (preferred) or AWS.
  • Operate and scale Kubernetes clusters (GKE preferred) and deploy services using Helm.
  • Manage containerized workloads using Docker.
  • Develop automation and internal tooling using Python to improve reliability and observability.
  • Integrate CI/CD pipelines with reliability and monitoring checks.
  • Mentor junior engineers, influence architectural decisions, and collaborate across engineering teams.


Required Qualifications:

  • 6+ years of experience as a DevOps Engineer, SRE, or related software engineering role, supporting production-grade systems.
  • Strong hands-on experience with cloud infrastructure on GCP (preferred) or AWS.
  • Proven expertise in operating Kubernetes-based platforms in production environments (GKE preferred).
  • Solid experience designing and maintaining highly available and resilient systems using SRE best practices.
  • Hands-on knowledge of SLIs, SLOs, error budgets, and reliability engineering principles.
  • Strong experience with observability and monitoring tools, including Prometheus, Grafana, Alertmanager, OpenTelemetry, and log platforms such as Loki / ELK / OpenSearch.
  • Demonstrated experience in incident management, on-call support, root cause analysis, and postmortems.
  • Proficiency in automation and tooling using Python, with additional scripting experience in Shell or Groovy.
  • Experience integrating CI/CD pipelines (Jenkins, GitHub) with deployment, monitoring, and reliability checks.
  • Strong understanding of microservices architectures, distributed systems, and containerized workloads.
  • Hands-on experience with Infrastructure as Code (IaC) tools such as Terraform or CloudFormation.
  • Good knowledge of cloud networking, security fundamentals, and access controls.
  • Strong analytical and problem-solving skills with a proactive operational mindset.
  • Excellent communication skills and the ability to collaborate effectively with cross-functional engineering teams.

What We Offer:

  • Bootstrapped and financially stable with high pre-money evaluation.
  • Above-industry remunerations.
  • Additional compensation tied to Renewal and Pilot Project Execution.
  • Additional lucrative business development compensation.
  • Firm-building opportunities that offer a stage for holistic professional development, growth, and branding.
  • An empathetic, excellence-driven, and result-oriented organization that believes in mentoring and growing teams with a constant emphasis on learning.