About this role
About the Role The Command Center Engineer is responsible for ensuring the seamless operation of critical IT infrastructure by providing 24/7 monitoring, incident management, and technical escalation support to maximize system uptime and minimize service disruptions. What You'll Do
- 24/7 Proactive Monitoring & Incident Response: Utilize BBPM, HP NNM, SolarWinds, Splunk, and Prometheus to detect and respond to system anomalies; conduct real-time log analysis and telemetry monitoring; implement event correlation to reduce alert noise and prioritize incidents.
- Escalation Management & Root Cause Analysis (RCA): Act as a key escalation point for L1 teams; collaborate with infrastructure, network, and application support; drive PIRs; assist in forensic analysis using SIEM and system logs.
- Operations Reporting & Metrics Analysis: Generate detailed operations reports with performance trends and MTTR; develop dashboards in ServiceNow, Power BI, or Grafana for real-time health insights; perform trend analysis for proactive improvements.
- Technical Process Management & Knowledge Base Development: Define/refine SOPs for incident handling and troubleshooting; maintain updated KB with runbooks; ensure ITIL-compliant Change, Incident, and Problem Management.
- Critical Incident Management & Crisis Coordination: Serve as Major Incident Manager during high-severity outages; coordinate war room calls and communications with stakeholders and executives; maintain Major Incident communications. What We're Looking For
- Proven experience in command center/onsite GCC operations with ability to work 24/7 on-call shifts.
- Deep technical expertise in IT infrastructure monitoring, incident management, and escalation.
- Hands-on experience with BBPM, HP NNM, SolarWinds, Splunk, Prometheus; strong log analysis and event correlation skills.
- Familiarity with SIEM, runbooks, and ITIL-based Change, Incident, and Problem Management.
- Proficiency in creating/consuming dashboards (ServiceNow, Power BI, Grafana) and generating MTTR metrics.
- Strong communication and stakeholder management; ability to lead war-room discussions and PIRs. Nice to Have
- Experience with forensic analysis leveraging SIEM and system logs.
- Prior Major Incident Management (MIM) experience and familiarity with post-incident reviews.
- Knowledge of ITIL best practices and hands-on problem-solving mindset. Compensation & Benefits
- Not disclosed