About this role
Job title: Site Reliability Engineer - Observability & Incident Management
About the Role
Join Ripple’s Technical Operations team as an Engineering-first Site Reliability Engineer focused on Observability and Incident Management. You will do hands-on reliability work (instrumentation, alerting, Terraform, production troubleshooting) while coaching stream-aligned product teams to mature operational practices. The role spans across Azure (80%) and AWS (20%) in a predominantly Windows-based environment, and you will help shape an early-stage incident management program.
What You'll Do
-
Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS; write NRQL queries for troubleshooting and reporting.
-
Define and implement SLOs/SLIs and error budgets; coach teams on using them to balance feature velocity with reliability and clearly communicate system health to stakeholders.
-
Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.
-
Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.
-
Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation, distributed tracing, and effective dashboard patterns.
-
Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources; enforce IaC governance across teams.
-
Author and troubleshoot Azure DevOps pipelines; support deployment visibility, change tracking, and release hygiene as it relates to production reliability.
-
Incident Management: Administer Incident.
-
IO—alert routing, notification workflows, Slack and OpsGenie integration, and runbook management; expanding from the current baseline.
-
Build incident management foundations: postmortems, on-call rotation design, escalation policies, incident severity classification, and response playbooks.
-
Track and report MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement with engineering teams.
-
Respond to and debrief on production incidents—providing real-time troubleshooting support and facilitating structured post-incident reviews.
-
Cross-Functional Enablement: enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance; collaborate with the Subsystems Platform Team to translate needs into self-service capabilities; create documentation and training materials that outlast individual engagements.
-
What We're Looking For
-
Core SRE Experience: 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations; proven ability to do hands-on work while coaching and mentoring teams; comfortable switching between builder and consultant modes; experience in Agile/Scrum environments.
-
Observability & Incident Management Expertise — Required: Expert-level hands-on experience with New Relic and strong NRQL proficiency; deep understanding of structured logging, metrics collection, distributed tracing, and designing effective dashboards and alerts; expertise defining and implementing SLOs/SLIs and error budgets; hands-on experience with incident management platforms; experience designing incident response workflows, on-call rotations, escalation policies, and facilitating post-incident reviews.
-
Cloud and IaC Experience: Demonstrated ability to work across Azure and AWS; Terraform for provisioning observability resources; experience authoring and troubleshooting Azure DevOps pipelines.
-
Collaboration and Coaching: Strong collaboration skills with cross-functional teams; demonstrated ability to coach teams toward greater operational maturity and to translate needs into self-service capabilities.