Talent Apply
Log in
All jobs
T

Site Reliability Engineer

TradeStation
Heredia, Costa Rica (Remote)
Remote

About this role

Job title: Site Reliability Engineer

What We Are Looking For: Do you enjoy tackling interesting problems with other smart people? Have you ever wondered how a major player in the Financial Technology industry operates? At TradeStation, you'll enjoy working in a supportive environment with the latest in microservices and streaming technology — all from the comfort of your home office.

We are seeking a Site Reliability Engineer to join the Order Execution – Streaming (OXS) team. OXS builds low-latency, scalable, and extensible services that provide access to order execution back-office data — the streams and APIs that deliver orders, positions, and balances to TradeStation and Monex clients across web, desktop, mobile, and third-party APIs.

You will be paired with senior SREs and a defined readiness program rather than dropped into the deep end, and you will progress from shadowing to independent ownership of infrastructure, observability, and on-call responsibilities. We are looking for a well-rounded engineer with genuine curiosity about how distributed systems fail, and the discipline to work carefully in an environment where real customer funds are at stake.

What You'll Be Doing:

Reliability and Cloud Infrastructure

  • Understand, execute, and embody Site Reliability Engineering principles
  • Work and collaborate in building and maintaining AWS cloud infrastructure for the OXS development and quality engineering teams to utilize
  • Author and maintain infrastructure as code, with all AWS infrastructure created through Stacker and terraform, and checked into GitLab as configuration-as-code
  • Administer and maintain Kubernetes (EKS) clusters and the OXS workloads that run on them — node groups and upgrades, RBAC and namespace policy, resource requests/limits and autoscaling, ingress and service networking, and deployment manifests and Helm/pipeline configuration
  • Build and exercise cross-region disaster recovery for OXS services — multi-AZ and multi-region failover, backup and restore of state stores, replication of Kafka/MSK and configuration, and periodic DR testing against defined RTO and RPO targets
  • Assist with building the necessary guardrails to keep services operational and secure
  • Assist with building templates, automation, and tooling that accelerate development and reduce operational toil
  • Work in a DevOps environment, where development teams own both the development and operational responsibilities
  • Provide technical guidance to developers and less experienced engineers on cloud-native architecture, resilience patterns, and operational readiness

Observability and Service Health

  • Instrument OXS services with OpenTelemetry, build and maintain dashboards, metrics, and alerting
  • Help define and measure service level indicators and objectives for order, position, and balance streaming, including stage-level latency budgets
  • Improve alert quality — reduce noise, standardize severity-based routing, and close monitoring gaps surfaced by incidents
  • Analyze logs, traces, and Kafka consumer lag to identify degradation before customers notice it
  • Profile and tune performance across the AWS stack — EC2/EKS instance and node sizing, container CPU/memory tuning, JVM and runtime settings, MSK and ElastiCache throughput, network and load-balancer configuration — and validate improvements through load testing against latency budgets
  • Profile and tune performance across the AWS stack — EC2/EKS instance and node sizing, container CPU/memory tuning, JVM and runtime settings, MSK and ElastiCache throughput, network and load-balancer configuration — and validate improvements through load testing against latency budgets
  • Evaluate emerging technologies and recommend improvements to platform reliability, scalability, and operational efficiency

Environments, Pipelines, and Release Support

  • Support and maintain the OXS environment state across environments
  • Maintain GitLab CI pipelines, container images, and ECR lifecycle configuration; keep builds and deployments repeatable, versioned, automated, and rollback-capable
  • Support release operations alongside developers and SDETs, including change request preparation and post-deployment validation
  • Follow TradeStation change management across different environments up to production, so that changes carry a complete audit trail of approvals

Incident Response and Production Support

  • Participate in the OXS on-call rotation, progressing from shadowing to Secondary and then Primary as you complete the team's on-call readiness program
  • Review dashboards and alerts at the start of the trading day and confirm system readiness ahead of market open
  • Triage and respond to incidents using OXS runbooks, escalating early rather than late
  • Contribute to blameless postmortems, root cause analysis, and the preventive-measure work items that follow them
  • Author and improve runbooks, SOPs, and operational documentation so that knowledge does not live in one person's head

Security and Cost

  • Help remediate cloud security findings identified through Wiz in coordination with the security team
  • Support cloud cost visibility and optimization for OXS AWS accounts, flagging spend anomalies and helping right-size infrastructure

Skills

You Bring:

  • Knowledge of Kubernetes administration — cluster and node group operations, upgrades, RBAC, scheduling and resource management, autoscaling, and troubleshooting workloads and control-plane issues
  • Knowledge and proficiency in one or more modern general-purpose programming languages (e.g. Python, C#, Golang, Bash)
  • Knowledge of AWS cloud infrastructure and core services (EKS, EC2, S3, ECR, IAM, MSK, ElastiCache, RDS)
  • Experience building or modifying cloud infrastructure as code (e.g. Stacker, CloudFormation, Terraform)
  • Experience with Continuous Integration tools, GitLab CI preferred (e.g. GitLab CI, Azure DevOps, Jenkins)
  • Familiarity with observability tooling and concepts — metrics, logs, traces, dashboards, and alerting (e.g. OpenTelemetry, Grafana, Datadog, Prometheus)
  • Excellent written and verbal communication skills, with the ability to write clear runbooks and incident updates and to assist developers in building cloud native applications
  • Familiarity working in an Agile environment and demonstrated success with structured testing practices such as automated unit testing, integration testing, TDD, and continuous delivery
  • Willingness to participate in an on-call rotation supporting a production trading system, and to be responsive during US equity and futures market hours
  • Demonstrated working knowledge of a cloud provider, containers, and a CI/CD pipeline, whether from professional experience, internships, or substantial personal projects
  • Experience operating and administering Kubernetes in production, including cluster upgrades, capacity and autoscaling decisions, and troubleshooting workloads under load
  • Experience participating in disaster recovery exercises or regional failover events for a production system
  • Understanding of high-availability and disaster recovery design in AWS — multi-AZ and cross-region architectures, failover strategies, backup/restore, and RTO/RPO tradeoffs
  • Experience diagnosing and tuning performance in cloud environments — latency and throughput analysis, right-sizing, capacity planning, and load or stress testing
  • Experience with distributed and scalable cloud architectures and techniques preferred
  • Familiarity with event streaming and messaging systems, especially Apache Kafka or AWS MSK, preferred
  • Familiarity with Redis, or another distributed key/state store, preferred
  • Exposure to Microsoft SQL Server administration or migration work preferred
  • Experience managing Linux deployments in the cloud preferred
  • Experience securing cloud deployments preferred
  • Advanced understanding of network and internet technologies

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →