About this role
Job title: Infrastructure Operations Manager (Team Lead)
About the Role We are seeking an experienced Infrastructure Operations Manager (Team Lead) to take overall responsibility for the devices and infrastructure in our AI Data Center. You will ensure operational excellence, lead a technical team, and meet client-driven SLAs and KPIs for high-performance AI workloads.
What You'll Do
- Site & Infrastructure Management: Overall accountability for all devices, systems, and infrastructure at the data center; ensure 24/7 reliability and optimal performance for high-performance AI workloads.
- Power, Cooling & Environment: Monitor and manage power, cooling, and environmental conditions; address operational risks proactively.
- HPC & GPU Systems: Oversee installation, configuration, and maintenance of HPC and GPU systems.
- Team Leadership & Development: Lead and mentor a team of engineers, technicians, and support staff; plan shift schedules and ensure on-call coverage.
- Client & Vendor Relations: Serve as primary POC for clients; provide regular reporting on SLAs and KPIs; collaborate with vendors for procurement, repairs, and upgrades.
- Inventory & Resource Management: Manage spare parts inventory; track usage and coordinate procurement.
- Technical Expertise: Maintain strong working knowledge of HPC, GPU installations, hardware/software solutions; troubleshoot issues and stay updated on trends in AI/HPC infrastructure.
- Operational Monitoring & Reporting: Implement and manage monitoring tools; prepare reports on performance, downtime, and efficiency; propose improvements.
- Travel & On-Call: On-call duties and occasional travel to other sites as business needs arise.
What We're Looking For
- Bachelor’s degree in Computer Science, Engineering, or a related field.
- 5+ years of experience managing data centers, particularly in HPC and GPU environments.
- Proven leadership skills with experience managing and developing teams.
- Technical expertise in HPC, GPU deployment, and associated hardware/software solutions.
- Strong understanding of power, cooling, and environmental systems in data centers.
- Excellent client-facing communication and reporting skills.
- Experience with inventory and spare parts management.
Nice to Have
- Certifications in data center management (e.g., CDCP, CDCS, or similar).
- Hands-on experience with NVIDIA GPUs, CUDA, and AI frameworks.
- Familiarity with hybrid cloud/HPC environments.
Compensation & Benefits
- Highly competitive package (base + equity) with reviews every 12 months.
- Human-First Flexibility: Flexible workplace with autonomy to shape your day around life's moments.
- Growth opportunities and dynamic progression plans; exposure to cutting-edge AI technologies.