Talent Apply
Log in
All jobs
R

Senior Site Reliability Engineer, Observability & Incident Management

Ripple

About this role

Job title: Senior Site Reliability Engineer, Observability & Incident Management

About the Role

This is an engineering-first role on Ripple's Technical Operations team focused on observability and reliability engineering. You will design instrumentation, build alert configurations, and write Terraform for monitoring resources across Azure and AWS, while coaching product teams to raise operational maturity and reliability.

What You'll Do

  • Design and implement monitoring, alerting, and dashboards in New Relic across Azure and AWS; write NRQL queries for troubleshooting and reporting.

  • Define and implement SLOs/SLIs and error budgets; coach teams on balancing feature velocity with reliability.

  • Lead alert noise reduction and signal quality engineering—tune thresholds, eliminate false positives, and ensure every alert is actionable.

  • Optimize observability costs through log ingestion management, pipeline rules, and New Relic configuration governance.

  • Partner with engineering teams to improve observability maturity: structured logging, metrics instrumentation, distributed tracing, and effective dashboard patterns.

  • Develop Terraform infrastructure as code for provisioning and managing monitoring resources—primary engineering responsibility.

  • Establish and enforce IaC governance standards for observability across teams, providing a repeatable model for monitoring resources.

  • Author and troubleshoot Azure DevOps pipelines; support teams with deployment visibility and release hygiene as it relates to production reliability.

  • Administer Incident.IO: alert routing, notification workflows, Slack and OpsGenie integration, and runbook management.

  • Build incident management foundations: PIR/postmortem processes, on-call rotation design, escalation policies, incident severity classification, and structured post-incident reviews.

  • Track MTTR, MTTD, and incident frequency; identify trends and drive continuous improvement in partnership with engineering teams.

  • Respond to and debrief on production incidents—providing real-time troubleshooting support and facilitating structured post-incident reviews.

  • Enable stream-aligned engineering teams to adopt improved observability and incident management practices through workshops, consultation, and hands-on guidance.

  • Collaborate with the Subsystems Platform Team to translate common needs into self-service observability and incident management capabilities; build documentation and training materials.

  • What We're Looking For

  • 7+ years in Site Reliability Engineering, DevOps, or Platform Engineering with a strong focus on observability and production operations.

  • Proven ability to deliver hands-on engineering work while coaching and mentoring teams—comfortable switching between builder and consultant modes.

  • Experience working in Agile/Scrum environments and collaborating effectively with cross-functional teams.

  • Observability & Incident Management Expertise — Required: Expert-level hands-on experience with New Relic and strong NRQL proficiency for troubleshooting and analysis.

  • Deep understanding of structured logging, metrics collection, distributed tracing, and designing effective dashboards and alerts.

  • Expertise defining and implementing SLOs/SLIs and error budgets for reliability management.

  • Hands-on experience with incident management platforms.

  • Experience designing incident response workflows, on-call rotations, escalation policies, and facilitating post-incident reviews.

  • Ability to work across Azure and AWS, with Windows-based infrastructure experience.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →