Talent Apply
Log in
All jobs
R

Observability Engineer

Railway

About this role

About the Role This role focuses on building and operating Railway's observability platform, enabling real-time ingestion and analysis of logs, metrics, and telemetry. You will design scalable, fault-tolerant systems that keep the platform fast, reliable, and self-managing, with an emphasis on immutable infrastructure and service-to-service communication. What You'll Do

  • Build and operate ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry
  • Build scalable, fault-tolerant alerting engines for notifying users, in real-time, of threshold breaches
  • Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application
  • Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams
  • Extend and build Golang/Rust GRPC services capable of supporting millions of users, and the tens of millions to come
  • Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible
  • Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success
  • Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption
  • Today this role leans heavily operational — keeping our existing observability stack fast, reliable, and scaling under load — with focused stretches of service-building work as we extend the platform What We're Looking For
  • A strong understanding of distributed systems. You enjoy building and operating fault tolerant, resilient, and scalable services
  • Experience operating and extending Observability stacks (Clickhouse, VictoriaMetrics, etc.) at scale
  • A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo
  • The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around
  • A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup
  • A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
  • A great set of communication skills

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →