About this role
Job title: Site Reliability Engineer
About the Role Site Reliability Engineer at Valtech who is passionate about experience innovation, observability, and reliability across platforms. You will bring 6+ years of experience and a growth mindset to push the boundaries of what's possible in high-impact, customer-facing systems, collaborating across global teams.
What You'll Do
- Maintain and improve observability systems (monitoring, logging, alerting).
- Define, adjust, and maintain Service Level Objectives (SLOs).
- Participate in incident resolution and on-call rotations (max 1 week/month).
- Drive proactive reliability improvements across platforms.
- Collaborate with teams to analyze failure scenarios and implement mitigations.
- Create and maintain runbooks for incident response and prevention.
- Eliminate non-value-adding tasks through automation and process optimization.
What We're Looking For
- Education: Bachelor's or Master's degree in Computer Science, Engineering, or a related technical field.
- 2+ years in DevOps, SRE, or Support Engineering roles.
- Experience with incident management in high-traffic, public-facing platforms.
- Strong scripting skills (Python, Bash, or PowerShell).
- Familiarity with CI/CD tools: GitHub Actions, Azure DevOps, GitLab, Jenkins.
- Experience with monitoring/APM tools: Datadog, New Relic, Dynatrace, Prometheus, Grafana.
- Basic knowledge of serverless services in AWS, Azure, or GCP.
- Proficiency with Docker and containerized environments.
- Excellent English communication skills (B2+ level).
- Experience working in international, cross-cultural teams.
Nice to Have
- Experience with Ansible, Chef, or other Infrastructure as Code tools.
- Background in e-commerce or high-impact platforms.
- Familiarity with runbook creation and failure scenario analysis.
- Exposure to corporate environments and agile methodologies.
Compensation & Benefits Not disclosed in the posting.