About this role
About the Role This role focuses on building and operating Railway's observability platform, enabling real-time ingestion and analysis of logs, metrics, and telemetry. You will design scalable, fault-tolerant systems that keep the platform fast, reliable, and self-managing, with an emphasis on immutable infrastructure and service-to-service communication. What You'll Do
- Build and operate ingestion pipelines to consume 1M+ RPS streams of logs, metrics, and other telemetry
- Build scalable, fault-tolerant alerting engines for notifying users, in real-time, of threshold breaches
- Craft rich backend observability APIs, working with product to build amazing experiences for instantly grokking their application
- Provide APIs to access realtime log/metrics streams to be consumed by the Dashboard and Product Teams
- Extend and build Golang/Rust GRPC services capable of supporting millions of users, and the tens of millions to come
- Define infrastructure that can be torn down, failed over, and reconstituted from scratch using principle of immutable infrastructure using Terraform and Ansible
- Write Engineering Requirement Documents to take something from idea, to defined tasks, to implementation, to monitoring it’s success
- Interface with our TypeScript and GraphQL edge to expose your microservice APIs for both internal and potentially external consumption
- Today this role leans heavily operational — keeping our existing observability stack fast, reliable, and scaling under load — with focused stretches of service-building work as we extend the platform What We're Looking For
- A strong understanding of distributed systems. You enjoy building and operating fault tolerant, resilient, and scalable services
- Experience operating and extending Observability stacks (Clickhouse, VictoriaMetrics, etc.) at scale
- A solid intuition about how long your solutions will last. All systems age. In startups, we can hope for 2-3 orders of magnitude, or 12-18mo
- The tact to implement your solution, creator monitors for it’s error boundaries, and document any requirements for when you’re not around
- A great sense of direction and prioritization when it comes to dealing with the ambiguity of an early stage startup
- A sense of grit to dive into a problem, implement a solution, scale that solution, and replace it when needed
- A great set of communication skills