About this role
A Site Reliability Engineer (SRE) is sought to support the reliability, availability, and performance of business-critical systems. The role focuses on AWS cloud infrastructure, DevOps practices, and core SRE disciplines. You will work closely with development, platform, and operations teams to ensure systems are stable, scalable, and well monitored.
- Reliability and Operations: Maintain high availability, scalability, and performance of production systems. Identify and reduce operational toil through automation and process improvement. Support design and implementation of fault-tolerant and resilient systems.
- AWS and Cloud Operations: Manage and operate systems hosted on AWS (EC2, EKS/ECS, RDS, S3, Lambda, CloudWatch, IAM, VPC). Support cloud deployments and infrastructure changes following best practices. Assist with backup, disaster recovery, and resiliency planning.
- DevOps and Automation: Work with CI/CD pipelines and DevOps tools to support reliable deployments. Use Infrastructure as Code tools such as Terraform or CloudFormation. Automate repetitive tasks using scripts (Python, Bash, etc.).
- Incident Management: Participate in production incident response, troubleshooting, and service restoration. Perform root cause analysis (RCA) and contribute to post-incident reviews. Help implement preventive actions to avoid incident recurrence.
- Observability: Configure and maintain monitoring, logging, and alerting using tools such as CloudWatch, Prometheus, Grafana, Splunk, or Dynatrace. Build dashboards to thoroughly track system health and reliability metrics. Improve alert quality to reduce noise and improve response times.
- Collaboration: Work closely with application and engineering teams to embed reliability into system design. Act as a strong team player, sharing knowledge and supporting team goals. Communicate effectively with technical and non-technical stakeholders.