About this role
About the Role The Senior Data Center Operations Engineer plays a critical, hands-on role in supporting the build-out and long-term operation of a high-performance, enterprise-scale data center environment supporting advanced compute and large-scale infrastructure deployments. This role requires experienced hardware, Linux systems, and data center operations expertise in highly available, production-critical environments.
What You'll Do
- Advanced hardware troubleshooting and repairs across server platforms (motherboards, CPUs, memory, storage), diagnosing failures and performing component-level replacements to minimize downtime.
- Conduct break/fix processes and root-cause analyses (RCA); identify failure trends and propose preventive improvements; contribute to tooling, automation, and process enhancements.
- Administer and support Linux systems and platforms, using Linux CLI tools for monitoring, diagnostics, and troubleshooting; deploy and configure servers across distributions (RHEL, Ubuntu, etc.); diagnose boot and OS issues in production; collaborate with engineering to resolve complex hardware/software problems.
- Data center operations: assist with equipment installation, structured cabling, and infrastructure validation; maintain accurate inventories; document repairs, changes, and configurations in ITSM/DCIM tools; ensure adherence to security, safety, and operations standards; serve as the principal escalation point for complex incidents; participate in on-call rotation (24/7).
- Collaboration and mentoring: coach technicians on hardware troubleshooting and best practices; work with network, storage and infrastructure teams to resolve cross-functional issues; contribute to knowledge sharing, documentation, and operational excellence; participate in continuous improvement initiatives.
What We're Looking For
- English proficiency (oral and written) to communicate effectively with global cross-functional teams and review English-language technical documentation.
- Advanced expertise in server hardware architecture and component-level troubleshooting.
- Excellent Linux systems and CLI diagnostic tool skills; strong problem-solving and attention to detail.
- Good understanding of networking basics and infrastructure components.
- Experience in structured operating environments (SOP, SLA, ticketing systems); familiarity with ITSM/DCIM tools (ServiceNow, Jira or equivalent).
- Experience with structured cabling and fiber optic connectivity.
- Ability to perform in production-critical, high-pressure environments; strong organizational and documentation skills.
- 5+ years of data center or infrastructure operations experience; significant experience in hardware repair/troubleshooting; production Linux experience (minimum 2 years); experience with enterprise server platforms; RCA and complex problem-solving track record; experience with ticketing and operational processes; experience in data center deployment or upgrades (a plus).
- Certifications: CompTIA A+, Server+, or Linux+; LPI or equivalent; vendor-specific certifications (preferred).
- Physical requirements: capable of lifting up to 50 lbs; able to work in a controlled-temperature, moderate-noise environment; capable of performing physical tasks (standing, walking, bending, kneeling) for extended periods.
Compensation & Benefits
- Salary details not disclosed. Benefits and on-call responsibilities associated with a 24/7 data center operations role may apply.