Do you want to work in one of the world's largest hyperscale infrastructure environments
We are seeking a Data Centre Critical Facilities Engineer to support the monitoring, triage, and coordination of operational incidents across critical data centre infrastructure within large-scale hyperscale and colocation environments.
Responsibilities:
- Monitor critical facilities infrastructure using enterprise monitoring platforms including BMS, DCIM, and EPMS.
- Respond to alarms across power, cooling, environmental, fire, and life safety systems supporting hyperscale data centre environments.
- Investigate infrastructure alerts, assess operational impact, and determine the appropriate escalation path.
- Act as the first point of contact for critical facilities incidents affecting live production environments.
- Coordinate with onsite Critical Facilities Engineers, vendors, Smart Hands teams, and site operations throughout incident resolution.
- Monitor utility power systems, UPS, battery plants, generators, ATS, PDUs, CRAH/CRAC units, chilled water systems, and environmental monitoring platforms.
- Manage incidents through to resolution, ensuring accurate documentation, timely communication, and operational updates.
- Track alarms, infrastructure events, and maintenance activities through JIRA and internal ticketing systems.
- Produce shift handover reports and maintain detailed operational documentation.
- Support planned maintenance and operational changes by following SOPs, MOPs, EOPs, and established runbooks.
- Participate in a 24x7 rotating shift supporting mission-critical data centre infrastructure.
Qualifications:
- 1-3+ years experience within Critical Facilities Operations, Data Centre Operations, Facilities Operations Centres (FOC), NOC, or mission-critical environments.
- Experience monitoring BMS, DCIM, EPMS, or similar infrastructure monitoring platforms.
- Understanding of critical power and cooling infrastructure within data centre environments.
- Familiarity with UPS systems, generators, ATS, PDUs, HVAC, CRAH/CRAC units, chilled water systems, or environmental monitoring.
- Experience working with incident management or ticketing systems such as JIRA, ServiceNow, or similar.
- Experience coordinating incidents across onsite engineering teams, vendors, and external service providers.
- Strong incident management, prioritisation, and operational coordination skills.
- Experience following SOPs, MOPs, EOPs, and operational runbooks.
- Excellent communication skills with the ability to work effectively in a fast-paced operational environment.