NextRaise Logo
NextRaise
JobsDiscover rolesJob TrackerTrack applied positionsMy ResumesBuild & optimize resumes
Tools
Job Match AnalyzerPaste JD, get fit scoreATS ScoreScan for ATS issues
Chrome Extension
ResumesJobsProfile
Jobs / Site Reliability Engineer in United States of America
3 hours agoBe an early applicant
Carislifesciences·3 hours ago
3 hours agoBe an early applicant

Site-Reliability Engineer, Application Operations

Irving, United States of AmericaMid · 2-5 years

Sign up free to see how well your resume matches this role.

About this role

At Caris, we understand that cancer is an ugly word—a word no one wants to hear, but one that connects us all. That’s why we’re not just transforming cancer care—we’re changing lives.

 

We introduced precision medicine to the world and built an industry around the idea that every patient deserves answers as unique as their DNA. Backed by cutting-edge molecular science and AI, we ask ourselves every day: “What would I do if this patient were my mom?” That question drives everything we do.

 

But our mission doesn’t stop with cancer. We're pushing the frontiers of medicine and leading a revolution in healthcare—driven by innovation, compassion, and purpose.

 

Join us in our mission to improve the human condition across multiple diseases. If you're passionate about meaningful work and want to be part of something bigger than yourself, Caris is where your impact begins.

Position Summary

Caris Life Sciences is one of the largest precision-oncology platforms in the world, serving hundreds of thousands of molecular cases a year and growing at double-digit rates. Behind every case is a matched molecular, imaging, and clinical-outcomes data estate few organizations anywhere can rival, and the clinical software that drives the lab's instruments and processes, captures results, and delivers each patient's report. When that software degrades, patient care waits; keeping it healthy is this role's mission.

The Site-Reliability Engineer, Application Operations, works on the dedicated Application Site-Reliability Engineering (App-SRE) team. The work is site-reliability engineering across the clinical application portfolio, with real production telemetry and a direct hand in shaping the run-operate discipline.

This is a production-operations craft role, distinct from application feature development, operating under governed privileged access and segregation-of-duties discipline. The engineer serves as a primary responder in the on-call rotation; executes runbooks and approved maintenance scripts with audit-grade discipline; implements application instrumentation, SLO monitors, and error-budget tracking; supports releases, deployment health verification, and rollback execution; and contributes to post-incident reviews. The team runs automation-first: recurring manual work is engineering backlog, and the engineer progressively automates away the toil they encounter rather than absorbing it.

Frontier AI coding assistants are standard-issue tooling, with agentic workflows spanning runbook authoring, alert-quality analysis, and operational automation. Characterization tests and golden-master replay validation serve as executable evidence for regulated change. Delivery runs on CI/CD with application-level observability, spanning application-performance monitoring and production telemetry, and risk-based release governance aligned with FDA Computer Software Assurance guidance. Modernization of established systems is active engineering work, not deferred maintenance.

Reporting to the Director, Application Site-Reliability Engineering, this practitioner-level individual contributor operates production clinical applications subject to SOX financial controls and FDA regulatory requirements, working safely and accurately under established, documented procedures.

Location: Irving, Texas (Dallas–Fort Worth), on-site/hybrid, minimum three days per week on campus; co-located with Caris's laboratory and clinical operations.

Job Responsibilities

  • Participate in a scheduled on-call rotation as a primary responder for production incidents across the SOX- and FDA-regulated clinical application portfolio; execute triage, escalation, and initial remediation steps in accordance with documented runbooks and incident-response playbooks.

  • Execute approved runbooks and maintenance scripts in the production environment; document all privileged actions in compliance with SOX ITGCs and FDA audit-trail requirements.

  • Implement and maintain application-layer instrumentation, including dashboards, alert thresholds, and SLO monitors, in coordination with the centrally operated observability platform.

  • Support production releases and deployments: coordinate deployment health verification, monitor application behavior after deployment, and execute rollback procedures when required.

  • Participate in post-incident reviews: contribute timeline reconstructions, identify contributing factors, and track remediation action items through to closure.

  • Maintain audit-ready production-access logs, privileged-action records, and role-change documentation to support SOX ITGC and FDA regulatory audits.

  • Automate recurring operational work: convert repeated manual interventions, diagnostics, and maintenance procedures into scripted, reviewed, pipeline-executed automation under the team's change controls.

  • Track the manual-intervention rate for assigned services and drive it down over time by retiring runbook steps into automation.

  • Author and update runbook entries and known-issue documentation as operational knowledge is gained; contribute to continuous improvement of the operate discipline.

  • Collaborate with application engineering teams to gather context during incidents and to validate fixes deployed to the production environment.

  • Monitor application SLO attainment and error-budget consumption; escalate proactively when budgets are at risk.

  • Contribute to the ongoing development of on-call tooling, alert quality, and operational dashboards within the application-layer observability framework.

  • Work AI-first: use AI coding assistants and agentic workflows as daily practice in monitoring, triage, runbook work, and automation scripting, with review as the quality gate.

Required Qualifications

  • Bachelor's degree in Computer Science, Information Systems, Software Engineering, or a closely related technical field, or equivalent practical experience.

  • 5+ years of professional experience in SRE, production operations, DevOps, or platform engineering in a production-support capacity.

  • 2+ years of direct, hands-on experience participating in an on-call rotation as a primary production responder for Tier 1 or business-critical systems.

  • Experience executing production runbooks, maintenance scripts, or change procedures in a production environment with documented privileged-action controls.

  • Experience implementing or maintaining application monitoring, alerting thresholds, or dashboards in a production observability platform.

  • Experience participating in post-incident review or blameless retrospective processes, including timeline reconstruction and corrective-action tracking.

  • Hands-on use of AI coding assistants for automation, scripting, or operational tooling.

Preferred Qualifications

  • Familiarity with SLO frameworks, error budgets, and associated alerting design patterns.

  • Experience reducing operational toil through scripting or automation in an SRE or production-operations setting.

  • Experience working in a SOX-controlled IT environment or a CLIA/CAP-regulated laboratory setting, including change-control ticketing, access-review processes, or audit-evidence collection.

  • Working knowledge of modern cloud-native observability at the application-instrumentation layer, including open standards for telemetry and tracing and application-performance-monitoring platforms.

  • Domain experience in clinical diagnostics, laboratory information systems, or digital health software.

  • Familiarity with SAST/DAST tooling and secure CI/CD pipelines, including pipeline-embedded security scanning.

  • Familiarity with deployment pipelines, container orchestration, and release-automation tooling from an operate and release-support perspective.

  • Certifications in cloud platforms or information-security disciplines relevant to production operations.

Physical Demands

  • Ability to sit, stand, and work at a computer for extended periods.

Training

  • All job-specific, safety, and compliance training is assigned based on the job functions associated with this employee.

Other

  • This role includes participation in a scheduled on-call rotation with required after-hours response to production incidents and critical service events, including evenings and weekends. Periodic travel may be required to support business needs and team on-sites.

Conditions of Employment:  Individual must successfully complete pre-employment process, which includes criminal background check, drug screening, credit check ( applicable for certain positions) and reference verification.

This job description reflects management’s assignment of essential functions. Nothing in this job description restricts management’s right to assign or reassign duties and responsibilities to this job at any time.

 

Caris Life Sciences is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, religion, color, national origin, gender, gender identity, sexual orientation, age, status as a protected veteran, among other things, or status as a qualified individual with disability.

H1B sponsor likely
AI tools
Apply faster with autofillThe NextRaise extension autofills your application in one click.Get the extension

Similar jobs

  • AI, Analytics & RPA Support SRE Manager at mizuhoNYC (1285), United States of America
  • System Reliability Engineer Program Manager at machHuntington Beach, United States of America
  • Senior Production Engineer - Northwest at Target301 STRANDER BLVD, United States of America
  • Director - Application Site-Reliability Engineering at carislifesciencesIrving, United States of America
  • Production Engineer (VCM) at westlakePlaquemine, United States of America
  • Production Engineer at sulzerLa Porte, United States of America