About this role
Job title: Site Reliability Engineer
What You'll Do
- Lead team technical discussions, especially around ongoing improvements in Reliability and Scalability
- Be involved in creating High Level Designs for new products and platforms
- Mentor junior SRE staff and enable them for success
- Lead incident response and post-mortem activities within your assigned service team
- Work with other Engineers in a cross-functional team to prioritise reliability improvements to address technical debt and toil
- Contribute to code to improve reliability
- Implement automation to reduce ongoing toil
What We're Looking For
- Minimum of 5+ years working experience in Software Development and/or Linux Systems Administration role.
- Strong interpersonal, written and verbal communication skills.
- Available to be scheduled in on-call rotation.
- Proficient as a Linux Production Systems Engineer, with experience managing large scale Web Services infrastructure.
- Development experience in one or more of the following programming languages: Python (preferred), Bash, Go, Java, C++, or Rust.
- Experience with at least 3 of the following topics: Distributed data storage at scale; NoSQL at scale; Data Aggregation technologies.