About this role
Job title: Senior Software Engineer- Site Reliability
About the Role As a Senior Site Reliability Engineer (SRE) at Kensho (S&P Global), you will be a hands-on technologist blending infrastructure expertise with software engineering to ensure the reliability, scalability, and security of production systems and customer facing services. You will own production systems end-to-end, design resilient architectures, and automate operations in a 24/7 on-call environment.
What You'll Do
- Own and operate production services supporting critical financial applications with focus on availability, performance, and reliability
- Design, build, and manage AWS infrastructure, including EKS clusters, across lower and production environments
- Provision and manage infrastructure using Terraform (Infrastructure as Code) with an automation-first mindset
- Deploy, scale, and troubleshoot applications on Kubernetes (including cluster creation, upgrades, lifecycle management)
- Build and maintain Python-based automation frameworks to reduce operational toil
- Monitor system health using metrics, logs, and alerts; tune alerts, dashboards, and runbooks
- Troubleshoot complex issues spanning clusters, networking, certificates, deployments, and application behavior
- Manage certificate lifecycle and expiration for secure service operation
- Collaborate with InfoSec, Vulnerability Management, and Network Security teams
- Collaborate with L1/L2 teams, leading on-call and incident response, root-cause analysis, and remediation
- Review new services for production readiness, resiliency, and secure design prior to release
- Establish production readiness standards, including deployment strategies, rollback plans, and observability
- Optimize infrastructure cost and resource utilization without compromising reliability
What We're Looking For
- 6+ years of experience in SRE, DevOps, Platform, or Infrastructure Engineering roles
- Strong software engineering background with hands-on Python for automation, tooling, and reliability
- Experience building or supporting scalable, distributed systems in production
- Deep experience with AWS, including IAM, networking, and access controls
- Hands-on with Kubernetes (EKS preferred): cluster creation, deployments, scaling, and troubleshooting
- Solid understanding of networking fundamentals (VPCs, routing, DNS, load balancing, security groups)
- Experience with CI/CD pipelines, deployment tools, and infrastructure automation
- Working knowledge of databases and query optimization, and understanding how applications behave under load
- Familiarity with Kafka or other messaging systems
- Comfortable conducting code reviews and participating in coding interviews
- Strong operational mindset with incident management and on-call rotations
- Clear communicator and collaborative teammate who values documentation and knowledge sharing
Nice to Have
- Demonstrated ownership of large-scale production systems
- Strong examples of Python-based automation or internal tooling
- Contributions to open-source projects, infrastructure platforms, or reliability tooling
- Experience working closely with security and compliance teams in regulated environments
- Technologies often used in this space such as AWS, EKS, Terraform, Jsonnet, Kubernetes, Helm, CI/CD tooling, Python, Prometheus, Grafana, PostgreSQL, Kafka, Linux
Compensation & Benefits Details not specified in the posting.