About this role
Job title: Site Reliability Engineer
What We Are Looking For: Do you enjoy tackling interesting problems with other smart people? Have you ever wondered how a major player in the Financial Technology industry operates? At TradeStation, you'll enjoy working in a supportive environment with the latest in microservices and streaming technology — all from the comfort of your home office.
We are seeking a Site Reliability Engineer to join the Order Execution – Streaming (OXS) team. OXS builds low-latency, scalable, and extensible services that provide access to order execution back-office data — the streams and APIs that deliver orders, positions, and balances to TradeStation and Monex clients across web, desktop, mobile, and third-party APIs.
You will be paired with senior SREs and a defined readiness program rather than dropped into the deep end, and you will progress from shadowing to independent ownership of infrastructure, observability, and on-call responsibilities. We are looking for a well-rounded engineer with genuine curiosity about how distributed systems fail, and the discipline to work carefully in an environment where real customer funds are at stake.
What You'll Be Doing:
Reliability and Cloud Infrastructure
- Understand, execute, and embody Site Reliability Engineering principles
- Work and collaborate in building and maintaining AWS cloud infrastructure for the OXS development and quality engineering teams to utilize
- Author and maintain infrastructure as code, with all AWS infrastructure created through Stacker and terraform, and checked into GitLab as configuration-as-code
- Administer and maintain Kubernetes (EKS) clusters and the OXS workloads that run on them — node groups and upgrades, RBAC and namespace policy, resource requests/limits and autoscaling, ingress and service networking, and deployment manifests and Helm/pipeline configuration
- Build and exercise cross-region disaster recovery for OXS services — multi-AZ and multi-region failover, backup and restore of state stores, replication of Kafka/MSK and configuration, and periodic DR testing against defined RTO and RPO targets
- Assist with building the necessary guardrails to keep services operational and secure
- Assist with building templates, automation, and tooling that accelerate development and reduce operational toil
- Work in a DevOps environment, where development teams own both the development and operational responsibilities
- Provide technical guidance to developers and less experienced engineers on cloud-native architecture, resilience patterns, and operational readiness
Observability and Service Health
- Instrument OXS services with OpenTelemetry, build and maintain dashboards, metrics, and alerting
- Help define and measure service level indicators and objectives for order, position, and balance streaming, including stage-level latency budgets
- Improve alert quality — reduce noise, standardize severity-based routing, and close monitoring gaps surfaced by incidents
- Analyze logs, traces, and Kafka consumer lag to identify degradation before customers notice it
- Profile and tune performance across the AWS stack — EC2/EKS instance and node sizing, container CPU/memory tuning, JVM and runtime settings, MSK and ElastiCache throughput, network and load-balancer configuration — and validate improvements through load testing against latency budgets
- Profile and tune performance across the AWS stack — EC2/EKS instance and node sizing, container CPU/memory tuning, JVM and runtime settings, MSK and ElastiCache throughput, network and load-balancer configuration — and validate improvements through load testing against latency budgets
- Evaluate emerging technologies and recommend improvements to platform reliability, scalability, and operational efficiency
Environments, Pipelines, and Release Support
- Support and maintain the OXS environment state across environments
- Maintain GitLab CI pipelines, container images, and ECR lifecycle configuration; keep builds and deployments repeatable, versioned, automated, and rollback-capable
- Support release operations alongside developers and SDETs, including change request preparation and post-deployment validation
- Follow TradeStation change management across different environments up to production, so that changes carry a complete audit trail of approvals
Incident Response and Production Support
- Participate in the OXS on-call rotation, progressing from shadowing to Secondary and then Primary as you complete the team's on-call readiness program
- Review dashboards and alerts at the start of the trading day and confirm system readiness ahead of market open
- Triage and respond to incidents using OXS runbooks, escalating early rather than late
- Contribute to blameless postmortems, root cause analysis, and the preventive-measure work items that follow them
- Author and improve runbooks, SOPs, and operational documentation so that knowledge does not live in one person's head
Security and Cost
- Help remediate cloud security findings identified through Wiz in coordination with the security team
- Support cloud cost visibility and optimization for OXS AWS accounts, flagging spend anomalies and helping right-size infrastructure
Skills
You Bring:
- Knowledge of Kubernetes administration — cluster and node group operations, upgrades, RBAC, scheduling and resource management, autoscaling, and troubleshooting workloads and control-plane issues
- Knowledge and proficiency in one or more modern general-purpose programming languages (e.g. Python, C#, Golang, Bash)
- Knowledge of AWS cloud infrastructure and core services (EKS, EC2, S3, ECR, IAM, MSK, ElastiCache, RDS)
- Experience building or modifying cloud infrastructure as code (e.g. Stacker, CloudFormation, Terraform)
- Experience with Continuous Integration tools, GitLab CI preferred (e.g. GitLab CI, Azure DevOps, Jenkins)
- Familiarity with observability tooling and concepts — metrics, logs, traces, dashboards, and alerting (e.g. OpenTelemetry, Grafana, Datadog, Prometheus)
- Excellent written and verbal communication skills, with the ability to write clear runbooks and incident updates and to assist developers in building cloud native applications
- Familiarity working in an Agile environment and demonstrated success with structured testing practices such as automated unit testing, integration testing, TDD, and continuous delivery
- Willingness to participate in an on-call rotation supporting a production trading system, and to be responsive during US equity and futures market hours
- Demonstrated working knowledge of a cloud provider, containers, and a CI/CD pipeline, whether from professional experience, internships, or substantial personal projects
- Experience operating and administering Kubernetes in production, including cluster upgrades, capacity and autoscaling decisions, and troubleshooting workloads under load
- Experience participating in disaster recovery exercises or regional failover events for a production system
- Understanding of high-availability and disaster recovery design in AWS — multi-AZ and cross-region architectures, failover strategies, backup/restore, and RTO/RPO tradeoffs
- Experience diagnosing and tuning performance in cloud environments — latency and throughput analysis, right-sizing, capacity planning, and load or stress testing
- Experience with distributed and scalable cloud architectures and techniques preferred
- Familiarity with event streaming and messaging systems, especially Apache Kafka or AWS MSK, preferred
- Familiarity with Redis, or another distributed key/state store, preferred
- Exposure to Microsoft SQL Server administration or migration work preferred
- Experience managing Linux deployments in the cloud preferred
- Experience securing cloud deployments preferred
- Advanced understanding of network and internet technologies