About this role
Job title: SRE - Network Infrastructure H/F/N
About the Role Site Reliability Engineer within the Infrastructure group at OVHcloud. You will help build a resilient, scalable and cost-efficient network platform, reduce operational friction, and automate durable solutions across our large-scale infrastructure.
What You'll Do
- Evaluate and prioritize incidents impacting OVHcloud's infrastructure and software platforms.
- Troubleshoot complex technical issues and coordinate cross-functional efforts to resolve them.
- Propose and implement best practices to ensure incidents are permanent fixes and do not recur.
- Participate in on-call rotations to ensure service continuity.
- Collaborate with development and infrastructure teams to remove bottlenecks, improve performance and reduce operational costs.
- Contribute to post-incident reviews and postmortems.
- Provide technical support to application owners and CI/CD pipeline stakeholders.
- Evolve and work in a network-focused IT environment.
What We're Looking For
- Solid Unix/Linux internals knowledge and confirmed networking skills.
- Strong software development experience (Python, Go, Perl, Bash).
- Proficiency with Infrastructure as Code tools (Ansible, Terraform).
- Experience operating distributed systems and knowledge of container technologies (Docker, Kubernetes).
- Strong understanding of CI/CD/CA tools.
- Autonomy, proactive attitude, and strong analytical mindset to resolve complex issues at scale.
- Practical experience with data pipelines and messaging/pub-sub systems (Redis, Kafka).
- Knowledge of monitoring tools (Prometheus, Grafana) and VXLAN.
Compensation & Benefits
- Hybride remote policy.
- Employee shareholding plan.
- Seniority recognition program.
- Vacation and sports subsidies.
- Company nursery and crèche (depending on site).
- Multicultural teams.
- Well-equipped premises.
- Online training and certification platform.
- Digitalized medical and social support for you and your family. You are trained on data up to October 2023.