Role Overview
We are seeking a seasoned Major Incident & Problem Manager to lead high-priority technology incident resolution, post-incident investigations, and operational service continuity across our core banking and infrastructure platforms. You will direct technical command bridges, establish clear operational accountability during crisis events, and ensure rapid service restoration in compliance with regulatory and ITIL standards.
Key Responsibilities
- Major Incident Command & Control: Take end-to-end ownership of critical incidents (Sev 1 / Sev 2), coordinating across cross-functional engineering, infrastructure, and application teams to minimize Mean Time to Restore (MTTR).
- Crisis Communication & Governance: Manage executive escalations, provide concise real-time situation updates to senior leadership and business stakeholders, and ensure full compliance with group technology standards.
- Problem Management & RCA: Facilitate post-incident reviews using structured analysis methodologies (e.g., 5 Whys, Fishbone) to identify underlying causes, eliminate repeat incidents, and track preventative actions to closure.
- Operational Reporting & Metrics: Monitor incident patterns, compile KPI dashboards (MTTR, SLA compliance, resolution timelines), and support audit and regulatory reporting deliverables.
- Continuous Service Improvement: Partner with Command Center and Infrastructure units to enhance automated alerting, runbook execution, and incident logging workflows.
Requirements
- Bachelor's degree in Computer Science or equivalent with around 8 years of relevant experience
- Proven track record leading Major Incident Management (MIM) and Problem Management within enterprise, high-availability IT environments.
- Strong working knowledge of ITIL service frameworks (ITIL certification required).
- Working familiarity with enterprise service management platforms (e.g., BMC Remedy, BMC Helix, ServiceNow).
- Broad technical literacy across enterprise environments: Application Support, End-of-Day (EOD) batch scheduling, infrastructure components (Linux/Unix, Storage, Network), middleware, and transactional workflows (e.g., payment channels).