NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in India
4 days ago
Apply with autofill
Apply with autofill
Nextgen·4 days ago
4 days ago

Sr. Observability Engineer II

Work From Anywhere-IndiaRemoteMid · 5-8 yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at nextgen

How you compare FREE

?
Your scoreYour score: not yet known
→
76
Top 10%Top 10%: 76 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in India.

Must-have skills for this role

  • dynatrace
  • dql
  • apms
  • observability

PDF or DOCX · no account needed

Apply faster with autofill FREEThe NextRaise extension autofills your application in one click.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Engineer and maintain the Dynatrace multi-tenant environment configuration, including management zones, alerting profiles, metric and custom events, tagging strategies, and access boundaries.
  • Own the end-to-end alert lifecycle, defining ownership, naming standards, severity mapping, alert disposition taxonomies, and production-readiness criteria.
  • Systematically eliminate alert fatigue across top problem patterns utilizing advanced tuning, thresholds, correlation, suppression, and alert retirement, while maintaining robust coverage for real, client-impacting issues.
  • Design and implement synthetic monitors and business-centric service-health models for critical client journeys and access paths, ensuring degradation is detected in service and business terms rather than raw infrastructure metrics.
  • Author and standardize Dynatrace Query Language (DQL) queries, notebooks, and dashboards for operations, leadership, and reliability reporting, establishing best-practice development standards for the team.
  • Collaborate with tools administrators to manage observability configurations using configuration-as-code practices, ensuring monitoring setups are versioned, peer-reviewed, and reproducible.
  • Partner with Tools, Automation and SRE teams to enhance alert-to-ticket enrichment (such as Salesforce ITSM integrations) with context and probable cause, minimizing triage times and enabling automated remediation.
  • Partner with SQL, platform, OS, and API specialists to resolve monitoring blind spots, ensure domain-specific alerts are meaningful, and verify that alerts are backed by robust runbooks.
  • Support service-level (P1–P4) measurements by establishing reporting baselines for corporate accountability, ensuring noise-reduction changes are paired with strict operational guardrails.
  • Maintain comprehensive reference materials, standard operating procedures, and observability standards documentation in Confluence to support continuous team enablement and operational onboarding.
  • Perform other duties that support the overall objective of the position.

What they're looking for

  • Bachelor’s Degree in Computer Science, Information Technology, or a related technical field. Or, any combination of education and experience which would provide the required qualifications for the position.
  • 7+ years of experience in observability, application performance monitoring (APM), monitoring engineering, site reliability engineering (SRE), or a closely related technology infrastructure field.
  • Demonstrable hands-on expertise with the Dynatrace platform, including Davis AI, DQL, dashboards, notebooks, management zones, alerting profiles, metric/custom events, and synthetic monitors.
  • Proven track record of designing, configuring, and tuning alerting systems at enterprise scale to successfully reduce noise while maintaining robust detection capabilities.
  • Experience integrating monitoring systems with enterprise ITSM/ticketing systems (such as Salesforce) for automated ticket routing and enrichment.
  • Experience working within highly regulated hosting environments (e.g., healthcare, HIPAA/HITRUST, SOC 2, or ISO 27001).

Nice to have

  • Dynatrace certification (or commitment to obtain certification within 1 year of hire).

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

Job Description:

The Sr. Observability Engineer II is a specialized engineering role serving as the subject matter expert for the Dynatrace observability platform. The primary purpose of this role is to mature the firm’s proactive monitoring capabilities by focusing on signal quality engineering and operational noise reduction. This position functions as a key technical liaison between Hosting Operations, Site Reliability Engineering (SRE), and other engineering teams to ensure the stability and reliability of our client-facing healthcare SaaS environment.

  • Engineer and maintain the Dynatrace multi-tenant environment configuration, including management zones, alerting profiles, metric and custom events, tagging strategies, and access boundaries.
  • Own the end-to-end alert lifecycle, defining ownership, naming standards, severity mapping, alert disposition taxonomies, and production-readiness criteria.
  • Systematically eliminate alert fatigue across top problem patterns utilizing advanced tuning, thresholds, correlation, suppression, and alert retirement, while maintaining robust coverage for real, client-impacting issues.
  • Design and implement synthetic monitors and business-centric service-health models for critical client journeys and access paths, ensuring degradation is detected in service and business terms rather than raw infrastructure metrics.
  • Author and standardize Dynatrace Query Language (DQL) queries, notebooks, and dashboards for operations, leadership, and reliability reporting, establishing best-practice development standards for the team.
  • Collaborate with tools administrators to manage observability configurations using configuration-as-code practices, ensuring monitoring setups are versioned, peer-reviewed, and reproducible.
  • Partner with Tools, Automation and SRE teams to enhance alert-to-ticket enrichment (such as Salesforce ITSM integrations) with context and probable cause, minimizing triage times and enabling automated remediation.
  • Partner with SQL, platform, OS, and API specialists to resolve monitoring blind spots, ensure domain-specific alerts are meaningful, and verify that alerts are backed by robust runbooks.
  • Support service-level (P1–P4) measurements by establishing reporting baselines for corporate accountability, ensuring noise-reduction changes are paired with strict operational guardrails.
  • Maintain comprehensive reference materials, standard operating procedures, and observability standards documentation in Confluence to support continuous team enablement and operational onboarding.
  • Perform other duties that support the overall objective of the position.

Education Required:

  • Bachelor’s Degree in Computer Science, Information Technology, or a related technical field.
  • Or, any combination of education and experience which would provide the required qualifications for the position.

Experience Required:

  • 7+ years of experience in observability, application performance monitoring (APM), monitoring engineering, site reliability engineering (SRE), or a closely related technology infrastructure field.
  • Demonstrable hands-on expertise with the Dynatrace platform, including Davis AI, DQL, dashboards, notebooks, management zones, alerting profiles, metric/custom events, and synthetic monitors.
  • Proven track record of designing, configuring, and tuning alerting systems at enterprise scale to successfully reduce noise while maintaining robust detection capabilities.
  • Experience integrating monitoring systems with enterprise ITSM/ticketing systems (such as Salesforce) for automated ticket routing and enrichment.
  • Experience working within highly regulated hosting environments (e.g., healthcare, HIPAA/HITRUST, SOC 2, or ISO 27001).

License/Certification Required:

  • Dynatrace certification (or commitment to obtain certification within 1 year of hire).

Knowledge, Skills & Abilities:

  • Knowledge of: Strong working knowledge of AWS infrastructure, Windows and Linux operating systems, and SQL Server database concepts. Advanced understanding of cloud-native observability frameworks, APM tools (specifically Dynatrace), and alert lifecycle management. Strong command of cloud computing (AWS), systems integration, database operations, and scripting (e.g., Python or PowerShell) to support automation.
  • Skill in: Excellent problem-solving capabilities. Strong written and verbal communication skills. Highly organized, detail-oriented, and self-driven, with a strong ownership mindset toward maintaining signal quality, reducing operational toil, and defending client service-level agreements.
  • Ability to: Proven ability to translate complex technical infrastructure metrics into direct service and business impacts. Ability to collaborate effectively across operations command analysts, SRE partners, database specialists, and leadership.

The company has reviewed this job description to ensure that essential functions and basic duties have been included. It is intended to provide guidelines for job expectations and the employee's ability to perform the position described. It is not intended to be construed as an exhaustive list of all functions, responsibilities, skills and abilities. Additional functions and requirements may be assigned by supervisors as deemed appropriate. This document does not represent a contract of employment, and the company reserves the right to change this job description and/or assign tasks for the employee to perform, as the company may deem appropriate.

NextGen Healthcare is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Company

Nextgen
Work From Anywhere-India

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Nextgen's careers site·first seen 16 Sept 2026·last verified 16 Sept 2026·How we source jobs

Similar jobs

  • Senior Site Reliability Engineer, Production Engineering at NVIDIABengaluru, India–match not yet calculated
  • Senior Mainframe Systems Programmer - Site Reliability Engineering at ensonoPune, India–match not yet calculated
  • Senior Site Reliability Engineer at Weekday AIBengaluru, India–match not yet calculated
  • Site Reliability Engineer III - Python, Grafana, Splunk, AWS, Jenkins at JPMorgan ChaseBengaluru, India–match not yet calculated
  • Site Reliability Engineer at AutodeskPune, India–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in India
  • DevOps Engineer jobs in India
  • Cloud Engineer jobs in India
  • Platform Engineer jobs in India
  • Site Reliability Engineer jobs in United States
  • Site Reliability Engineer jobs in United Kingdom