About this role
Job Title: Project Manager (Major Incident Management & NOC)
Location: Onsite – Wilmington, DE (Day1 Onsite)
Employment Type: Full-time
Experience: 10+ years in IT Operations / NOC / Major Incident Management
Role Summary
The Project Manager is responsible for Major Incident Management & NOC Teams.
This role leads the team for MIM and NOC functions who drive Major Incident (P1/P2) execution, ensures rapid service restoration, and continuously improves operational maturity through problem management, automation, observability enhancements, and SLA governance.
The role requires a mix of strong incident leadership, technical depth across infrastructure and applications, and people/process management to ensure stability, availability, and performance across critical services.
Key Responsibilities
-
A) Manage Team of Major Incident Managers (Command & Control)
-
Own the Major Incident (P1/P2) process from detection to resolution, including war-room leadership, stakeholder updates, and closure.
-
Ensure structured triage, containment, workaround, and restoration.
-
Drive cross-functional coordination (App, Infra, Network, Security, DB, Cloud, Vendor teams) to reduce MTTR.
-
Ensure high-quality incident communications: executive summaries, impact analysis, ETAs, customer/business comms.
-
Lead and facilitate Post Incident Reviews (PIR/RCA); ensure actionable corrective/preventive actions (CAPA).
-
Identify recurring issues and trigger Problem Management with measurable reduction plans.
-
ITIL v4 Foundation (preferred).
-
Reduced MTTD and MTTR for P1/P2 incidents.
-
Improved SLA compliance and reduction in escalation breaches.
-
Reduced repeat incidents via problem management and preventive actions.
-
Improved alert quality: lower false positives, better signal-to-noise ratio.
-
Strong PIR/RCA compliance: on-time RCAs with measurable preventive outcomes.
-
Improved NOC operational maturity: SOP adherence, shift handover quality, audit readiness.
-
B) NOC Leadership & Operations
-
Manage the NOC team responsible for 24x7 monitoring, alert triage, event correlation, escalation, and ticket quality.
-
Establish/maintain standard operating procedures (SOPs), runbooks, escalation matrices, and on-call models.
-
Ensure NOC meets SLAs/OLAs, improves alert fidelity, and reduces noise through tuning and automation.
-
Manage handover governance between shifts; maintain service continuity and operational hygiene.
-
C) Service Reliability & Continuous Improvement
-
Drive operational improvements: monitoring coverage, SLO/SLA alignment, incident prevention, and resiliency initiatives.
-
Partner with Engineering/Platform teams on observability strategy, proactive detection, and reliability patterns.
-
Track and report operational metrics: MTTD, MTTR, incident volume, re-open rate, SLA compliance, and trends.
-
Support readiness for audits and compliance: evidence collection, process adherence, and risk mitigation.
-
D) Stakeholder & Vendor Management
-
Interface with business units and vendors as appropriate to coordinate incident response and service delivery.