Talent Apply
Log in
All jobs
R

Site Reliability Engineer - Observability & Incident Management

Ripple

About this role

Job title: Site Reliability Engineer - Observability & Incident Management

About the Role

Join Ripple’s Technical Operations team as an Engineering-first Site Reliability Engineer focused on Observability and Incident Management. You will do hands-on reliability work (instrumentation, alerting, Terraform, production troubleshooting) while coaching stream-aligned product teams to mature operational practices. The role spans across Azure (80%) and AWS (20%) in a predominantly Windows-based environment, and you will help shape an early-stage incident management program.

What You'll Do

  • Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS; write NRQL queries for troubleshooting and reporting.

  • Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and clearly communicate system health to stakeholders.

  • Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.

  • Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.

  • Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation, distributed tracing, and effective dashboard patterns.

  • Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources; enforce IaC governance across teams.

  • Author and troubleshoot Azure DevOps pipelines; support deployment visibility, change tracking, and release hygiene as it relates to production reliability.

  • Incident Management: Administer Incident.

  • IO—alert routing, notification workflows, Slack and OpsGenie integration, and runbook management; expanding from the current baseline.

  • Build incident management foundations: postmortems, on-call rotation design, escalation policies, incident severity classification, and response playbooks.

  • Track and report MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement with engineering teams.

  • Respond to and debrief on production incidents—providing real-time troubleshooting support and facilitating structured post-incident reviews.

  • Cross-Functional Enablement: enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance; collaborate with the Subsystems Platform Team to translate needs into self-service capabilities; create documentation and training materials that outlast individual engagements.

  • What We're Looking For

  • Core SRE Experience: 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations; proven ability to do hands-on work while coaching and mentoring teams; comfortable switching between builder and consultant modes; experience in Agile/Scrum environments.

  • Observability & Incident Management Expertise — Required: Expert-level hands-on experience with New Relic and strong NRQL proficiency; deep understanding of structured logging, metrics collection, distributed tracing, and designing effective dashboards and alerts; expertise defining and implementing SLOs/SLIs and error budgets; hands-on experience with incident management platforms; experience designing incident response workflows, on-call rotations, escalation policies, and facilitating post-incident reviews.

  • Cloud and IaC Experience: Demonstrated ability to work across Azure and AWS; Terraform for provisioning observability resources; experience authoring and troubleshooting Azure DevOps pipelines.

  • Collaboration and Coaching: Strong collaboration skills with cross-functional teams; demonstrated ability to coach teams toward greater operational maturity and to translate needs into self-service capabilities.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →