Talent Apply
Log in
All jobs
A

Site Reliability Engineer

Avid
Remote, Philippines
Remote

About this role

Job title: Site Reliability Engineer

About the Role Site Reliability Engineer (Remote, Philippines) responsible for ensuring the reliability, performance, and scalability of our cloud infrastructure and production systems. You will collaborate with cross-functional engineering teams to design resilient architectures, automate deployments, and deliver a highly available platform.

What You'll Do

  • Champion and continuously improve platform reliability, observability, and DevOps culture across the engineering organization.
  • Define and track SLAs, SLOs, and SLIs to drive reliability goals and monitor service health across the platform.
  • Design, implement, and tune application and component monitoring, alerting and dashboards using Prometheus, Grafana, CloudWatch and Elastic (or similar tools).
  • Improve and harden core systems in conjunction with the larger cloud engineering team, including: designing, operating, and optimizing Kubernetes workloads on Amazon EKS; implementing Istio service mesh; building and managing GitOps pipelines with ArgoCD; automating CI/CD with GitHub Actions; provisioning infrastructure with Terraform; and managing AWS services (RDS, OpenSearch, IAM).
  • Participate in a 24/7 on-call rotation, handle incident response, perform postmortems, and maintain up-to-date runbooks.
  • Secure applications and infrastructure with tools like Snyk and follow security best practices.
  • Manage edge and DNS configurations with Cloudflare and Route 53 to ensure performance and global availability.
  • Operate and tune AWS services such as RDS, OpenSearch, and IAM, supporting data and identity needs.

What We're Looking For

  • Bachelor’s degree in Information Technology, Computer Science, Software Engineering, or related fields.
  • 5+ years of experience in Site Reliability Engineering, DevOps, or equivalent.
  • Strong proficiency with Kubernetes (preferably Amazon EKS) and containerized application deployments.
  • Proficiency with observability stacks (Prometheus, Grafana, ELK) and alerting best practices.
  • Strong scripting skills (Bash, Python, or similar) for automation and tooling.
  • Hands-on experience with infrastructure-as-code tools such as Terraform, and GitOps workflows using ArgoCD.
  • Experience with AWS services including RDS, IAM, OpenSearch, CloudWatch, and Route 53.
  • Experience with CI/CD automation using GitHub Actions or similar tools.
  • Knowledge of service mesh technologies such as Istio.
  • Understanding of security best practices in cloud environments, including vulnerability scanning and remediation.
  • Excellent troubleshooting skills, especially in distributed, cloud-native systems.
  • Strong communication skills and ability to work cross-functionally in a collaborative environment.

Nice to Have

  • Experience with security tooling and vulnerability management.
  • Additional cloud certifications or hands-on cloud architecture experience.
  • Experience with on-call incident management and postmortems.

Compensation & Benefits

  • Attractive benefits package including health & life insurance, referral rewards, and generous leave policies to ensure a healthy work-life balance.
  • Remote work model offering flexibility to balance work and life.
  • Access to development programs with strong support and mentoring to help you grow and advance within the company.

Your next opportunity starts here

Prepare, apply, track, interview and get hired — all from one platform, with AI in your corner.

Download app

Or sponsor Premium for someone who's job hunting →