NextRaiseNextRaiseFind jobs
Sign inSign up free
Jobs / Site Reliability Engineer in Australia
4 days ago
Apply with autofill
Apply with autofill
Preciselyinternationaljobs·4 days ago
4 days ago

Site Reliability Engineer

AustraliaHybridMid · 3+ yearsSite Reliability Engineer

Sign up free to see how well your resume matches this role.

Boost your chances at preciselyinternationaljobs

How you compare FREE

?
Your scoreYour score: not yet known
→
39
Top 10%Top 10%: 39 out of 100

Top 10% of NextRaise users matched against Site Reliability Engineer roles in Australia.

Must-have skills for this role

  • terraform
  • aws
  • linux
  • ansible

PDF or DOCX · no account needed

Apply faster with autofill FREEpreciselyinternationaljobs uses Greenhouse - autofill it instead of retyping.careers.example.com/applyAutofillingFull namePriya SharmaEmailpriya.sharma@example.comPhone+49 30 1234567LocationBerlGet the extension

What you'll do

  • Set and maintain reliability standards across CCX, RCX, and HMS, including SLOs, SLIs, error budgets, alerting, logging, tracing, and MTTR improvements.
  • Build and maintain infrastructure-as-code, deployment automation, monitoring, and observability tooling using Terraform, Ansible, Datadog, Python, and Bash.
  • Guide CI/CD and deployment standards while partnering with engineering teams to embed reliability, scalability, backup, recovery, and failure-mode planning into service design.
  • Review designs and lead Operational Readiness Reviews to validate production and disaster recovery readiness.
  • Lead major incident response, including incident command and clear stakeholder communication.
  • Produce root cause analyses, identify recurring issues, and implement preventive automation to reduce manual effort and improve reliability.
  • Maintain runbooks, reliability backlogs, and shared knowledge on monitoring gaps, incidents, lessons learned, and operational risks.
  • Use Precisely-provided AI tools for infrastructure code, incident analysis, troubleshooting, testing, runbooks, and documentation.
  • Ensure infrastructure meets security, compliance, vulnerability-remediation, and data-protection requirements, including applicable SOC 2 and FedRAMP standards.
  • Coach engineers, contribute to cross-team reviews, stay current with SRE practices, and participate in the rotating on-call schedule for critical escalations and changes.

What they're looking for

  • Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
  • 3+ years of systems or infrastructure engineering experience in an enterprise production environment.
  • Strong proficiency with Linux (RHEL/Oracle Linux) in a multi-site, multi-environment context.
  • Hands-on experience with infrastructure-as-code tools: Terraform and/or Ansible.
  • Experience deploying and managing workloads in AWS (EC2, ECS, S3, VPC, IAM).
  • Proficiency with at least one scripting language (Python, Bash) for automation development.
  • Demonstrated experience building and maintaining monitoring and alerting systems (Datadog preferred).
  • Solid understanding of TCP/IP networking, DNS, load balancing, and distributed systems.
  • Experience with CI/CD pipeline design and deployment automation standards.
  • Strong analytical skills; ability to perform structured root cause analysis and post-incident review.
  • Ability to define SLOs and lead Operational Readiness Reviews (ORRs); comfortable partnering with engineering teams on production readiness.
  • Demonstrated ability to work cross-functionally with engineering teams on reliability standards and observability requirements.

Nice to have

  • Experience with containerization and orchestration (Docker, ECS, Kubernetes).
  • Familiarity with GitOps workflows and source control best practices (Git, GitLab).
  • Knowledge of enterprise virtualization platforms in a hybrid cloud context.
  • Understanding of change management and ITIL operational practices.
  • Experience with enterprise security tooling (Qualys, CrowdStrike, Rapid7).
  • AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.

Summarised by NextRaise from the employer’s description, which follows in full below.

Full description from employer

At Precisely, we're not just building software — we're shaping the future of data integrity. As a global leader in data quality, data enrichment, and location intelligence, Precisely helps thousands of the world's most trusted brands make confident decisions with data they can rely on. We're an AI-first organization, which means artificial intelligence isn't a buzzword here — it's woven into how we build products, how we work, and how we think about solving complex problems for our customers. When you join Precisely, you join a team of curious, driven innovators who believe that better data makes the world run better. If you're ready to do meaningful work at the intersection of AI and data — and help define what's possible — we'd love for you to apply!

 

Overview:

The Site Reliability Engineer (SRE) is responsible for the reliability, performance, and scalability of Precisely's infrastructure platforms across CEDAR (CCX) — an on-premises, private cloud managed services environment; RapidCX (RCX) — an AWS cloud environment for SaaS-delivered customer communications management; and Hosted Managed Services (HMS) — an AWS cloud environment supporting managed client deployments.

This role bridges software engineering and systems operations, building automation, observability tooling, and reliability standards to ensure platform availability and operational excellence. SREs are enabling partners: they set reliability standards, define what 'reliable' looks like for each service, validate production readiness, and coach engineering teams on operational best practices. Engineering teams own the reliability outcomes of the services they build; the SRE ensures they have the standards, tooling, and guidance to meet them.

What you will do:

  • Set and maintain reliability standards across CCX, RCX, and HMS, including SLOs, SLIs, error budgets, alerting, logging, tracing, and MTTR improvements. 
  • Build and maintain infrastructure-as-code, deployment automation, monitoring, and observability tooling using Terraform, Ansible, Datadog, Python, and Bash. 
  • Guide CI/CD and deployment standards while partnering with engineering teams to embed reliability, scalability, backup, recovery, and failure-mode planning into service design. 
  • Review designs and lead Operational Readiness Reviews to validate production and disaster recovery readiness. 
  • Lead major incident response, including incident command and clear stakeholder communication. 
  • Produce root cause analyses, identify recurring issues, and implement preventive automation to reduce manual effort and improve reliability. 
  • Maintain runbooks, reliability backlogs, and shared knowledge on monitoring gaps, incidents, lessons learned, and operational risks. 
  • Use Precisely-provided AI tools for infrastructure code, incident analysis, troubleshooting, testing, runbooks, and documentation. 
  • Ensure infrastructure meets security, compliance, vulnerability-remediation, and data-protection requirements, including applicable SOC 2 and FedRAMP standards. 
  • Coach engineers, contribute to cross-team reviews, stay current with SRE practices, and participate in the rotating on-call schedule for critical escalations and changes. 

What we are looking for:

Required:

  • Educational requirements (equivalent work experience will be accepted in place of the education requirement): Bachelor's degree in Computer Science, Information Systems, Engineering, or equivalent practical experience.
  • Years of experience: 3+ years of systems or infrastructure engineering experience in an enterprise production environment.
  • Years of experience needed with specific skills: Not separately specified beyond the overall experience requirement above.
  • Specific technical or software skills required:
  • Strong proficiency with Linux (RHEL/Oracle Linux) in a multi-site, multi-environment context.
  • Hands-on experience with infrastructure-as-code tools: Terraform and/or Ansible.
  • Experience deploying and managing workloads in AWS (EC2, ECS, S3, VPC, IAM).
  • Proficiency with at least one scripting language (Python, Bash) for automation development.
  • Demonstrated experience building and maintaining monitoring and alerting systems (Datadog preferred).
  • Solid understanding of TCP/IP networking, DNS, load balancing, and distributed systems.
  • Experience with CI/CD pipeline design and deployment automation standards.
  • Strong analytical skills; ability to perform structured root cause analysis and post-incident review.
  • Ability to define SLOs and lead Operational Readiness Reviews (ORRs); comfortable partnering with engineering teams on production readiness.
  • Demonstrated ability to work cross-functionally with engineering teams on reliability standards and observability requirements.
  • Necessary certifications: None required (see Preferred Skills below).
  • Travel is required: No — approximately 0%.

AI Skills/Knowledge:

Active use of Precisely-provided AI tools (GitHub Copilot, Claude, or equivalent) for code development, troubleshooting, and documentation is a required baseline for this role, not a differentiator.

  • Apply AI tools for infrastructure-as-code development, incident analysis, and runbook authoring.
  • Use AI tools to accelerate troubleshooting and solution testing.
  • Maintain working fluency with Precisely-approved AI coding assistants as part of daily practice.

Preferred Skills (a plus but not required):

  • Experience with containerization and orchestration (Docker, ECS, Kubernetes).
  • Familiarity with GitOps workflows and source control best practices (Git, GitLab).
  • Knowledge of enterprise virtualization platforms in a hybrid cloud context.
  • Understanding of change management and ITIL operational practices.
  • Experience with enterprise security tooling (Qualys, CrowdStrike, Rapid7).
  • AWS Solutions Architect, SysOps Administrator, or DevOps Engineer certification.

 

#LI-KM1                                                                                    #LI-Remote

 

Application and Interview Impersonation Notice:  Impersonating another individual when applying for employment, and/or participating in an interview process to assist another individual in obtaining employment, with Precisely Software Incorporated (“Precisely”) is unlawful.  If Precisely identifies such fraudulent conduct, then as applicable and to the extent permitted by law, the application will be rejected, an offer (if made) will be rescinded, or the employment will be terminated, and legal action may be taken against the impersonators.

The personal data that you provide as a part of this job application will be handled in accordance with relevant laws. For more information about how Precisely handles the personal data of job applicants, please see the Precisely Candidate Privacy Notice

Company

Preciselyinternationaljobs
Australia

Company facts come from this company's own listings. We only show what the postings themselves carry.

Sourced from Preciselyinternationaljobs's careers site·first seen 17 Sept 2026·last verified 17 Sept 2026·How we source jobs

Similar jobs

  • Site Reliability Engineer at leidosCanberra, Australia–match not yet calculated
  • Senior Site Reliability Engineer - Observability at WestpacSydney, Australia–match not yet calculated
  • Reliability Engineer at RTXhenderson, Australia–match not yet calculated
  • Senior Platform Reliability Engineer at firmusMelbourne, Australia–match not yet calculated
  • Senior Storage Production Engineer - DGX Cloud at NVIDIAAustralia, Remote–match not yet calculated

Browse more jobs

  • Site Reliability Engineer jobs in Australia
  • Systems Engineer jobs in Australia
  • Platform Engineer jobs in Australia
  • Network Engineer jobs in Australia
  • Site Reliability Engineer jobs in United States
  • Site Reliability Engineer jobs in India