About this role
Job title: Site Reliability Engineer - Object Storage (F/H/N)
About the Role Join OVHcloud's Object Storage SRE team to build the most performant Object Storage offer and pilot IA integration in operations. In a cloud environment, you will design and evolve intelligent agents to automate incident detection/resolution, improve alert relevance, and manage preproduction testing, non-regression and operational performance, aligned with product and platform goals. As a Site Reliability Engineer in this department, you will help evolve, industrialize and maintain the operation of all our Object Storage products.
What You'll Do
- Use and integrate AI code assistants and AI agents in workflows to improve monitoring, alerting and incident detection on Object Storage platforms.
- Design and integrate intelligent agents capable of assisting or automating incident resolution workflows and continuous improvement.
- Contribute to reducing MTTD (Mean Time to Detection) and MTTR (Mean Time to Recovery) via automation driven by these agents and by procedures.
- Ensure high availability, reliability and security of Object Storage platforms; track performance indicators and contribute to their improvements.
- Ensure clients receive full technical support when needed and implement, apply and automate procedures to resolve common issues.
- Contribute to evolution of deployment, packaging, monitoring and alerting tools, with smooth integration of agents and AI tools into existing infrastructure and future projects.
- Challenge software/hardware architectures to improve performance, high availability and scalability.
- Monitor product adoption and customer usage, collaborating with technical and commercial teams to enrich backlog and roadmap.
- Write technical documentation and runbooks related to AI agents, automations and incident scenarios.
What We're Looking For
- Proficient in GNU/Linux administration.
- Experience in integrating/using AI agents (LLMs) in daily work.
- Mastery of one or more scripting languages (Python).
- Experience in automation and deployment (Puppet, Ansible).
- Experience working on complex microservices architectures.
- Proficient with monitoring/observability tools (Icinga / Prometheus / Alertmanager).
- Experience with large-scale infrastructure orchestration (Temporal).
- Nice to have: knowledge of AWS S3 API.
- Nice to have: experience handling large data volumes.
Nice to Have
- Knowledge of AWS S3 API
- Experience working with large data volumes
Compensation & Benefits
- Hybrid telework policy
- Employee stock ownership plan
- Seniority recognition program
- Vacation and sports subsidies
- Company nursery and on-site childcare (where available)
- Multicultural teams and well-equipped premises
- Online training and certification platform
- Digitalized medical and social support for you and your family
OVHcloud emphasizes diversity and inclusion, offering a collaborative environment with opportunities to grow and coconstruct the future of AI-assisted operations.