Head of Technical Delivery
Confidential- Posted 18 hours ago
- Be among the first 10 applicants
Job Description
Head of Technical Delivery, APACGPU Infrastructure Architecture & OperationsLocation: Singapore | Employment Type: Full-time | Seniority: Director level Travel: Approximately 30% within APAC | Reporting Line: Reports to APAC Leader
About Company
It is an enterprise GPU infrastructure provider enabling high-performance compute, artificial intelligence, and deep learning workloads across the globe. We design, deploy, and operate resilient, large-scale GPU clusters built on cutting-edge network fabrics, advanced storage architectures, and tier-one OEM partnerships. As we expand across the Asia-Pacific region, we are building a world-class engineering presence in Singapore to deliver mission-critical infrastructure to our enterprise and hyperscale clients.
Role Overview
It is seeking a seasoned technical leader to serve as the Head of Technical Delivery for the APAC region. In this role, you will hold ultimate accountability for technical execution across the entire lifecycle of our client deployments—spanning pre-freeze architecture validation, infrastructure delivery, rigorous acceptance testing, and long-term operational performance.
You must possess the technical authority to independently design complex GPU architectures from first principles. Whether evaluating customer-supplied blueprints or leading bespoke in-house designs, you will judge technical readiness, identify gaps, manage risk, and enforce stringent build and operational standards. In addition, you will establish our Singapore engineering hub, scale the regional team, and shape HydraHost's global engineering playbook.
Key Responsibilities
Architecture & BOM Ownership
- Own all cluster architecture and Bill of Materials (BOM) decisions across GPU compute systems, scale-up/scale-out fabrics, storage, and management layers.
- Rigorously evaluate customer-supplied designs against business requirements, OEM guidelines (e.g., NVIDIA), and industry standards; document risk profiles, technical deviations, and operational trade-offs.
- Lead engineering design processes for HydraHost-led deployments from initial specification through architectural freeze.
Technical Procurement & Supplier Qualification
- Direct technical bid evaluations, qualifying new hardware, nominated suppliers, and proposed component substitutions.
- Verify and formalize comprehensive vendor support coverage across OEMs, product lines, configurations, and service-level agreements (SLAs).
Quality Assurance & Acceptance Standards
- Define and execute regional acceptance criteria covering hardware health, network topology, link integrity, sustained stability, performance benchmarking, and failover/recovery.
- Leverage diagnostic frameworks and telemetry tools (including DCGM, NCCL, and HPL) to investigate system anomalies, withholding final sign-off until all acceptance criteria are met.
Configuration Management & Change Control
- Enforce baseline controls across firmware, drivers, operating systems, and network software with fully traceable as-built documentation.
- Institute formal change-management processes requiring peer-reviewed rollback plans; resolve complex compatibility issues through reproducible testing and direct OEM escalation.
Commissioning & Operational Handover
- Direct system commissioning, validation, and multi-party witness testing across integration boundaries.
- Resolve technical defects and ensure all facility, power, and environmental prerequisites are satisfied prior to operational mobilization.
Operational Readiness & Availability
- Design and execute the regional support, monitoring, spares management, hardware repair, and escalation models in partnership with central operations.
- Validate automated recovery mechanisms and lead major root-cause investigations to implement permanent corrective actions.
Leadership & Organizational Development
- Hire, mentor, and lead high-performing infrastructure engineers in Singapore in alignment with regional growth plans.
- Develop reusable engineering standards, automation tools, and runbooks, contributing regional insights to HydraHost's global engineering leadership.
- Act as the primary technical authority for HydraHost in customer and partner interactions, translating complex technical trade-offs into actionable business decisions.
Candidate Qualifications
- Experience: 12+ years of hands-on experience in GPU, High-Performance Computing (HPC), or cloud infrastructure engineering, with at least 5+ years in a technical leadership capacity.
- Architectural Depth: Proven track record of personally designing, building, and operating production GPU cluster architectures and BOMs.
- Network & GPU Topology: Expertise in NVIDIA GPU-system architectures, GPU-to-NIC topologies, and high-performance interconnects (RoCE and/or InfiniBand), with strong judgment regarding interoperability and redundancy.
- Ecosystem Knowledge: In-depth understanding of the server OEM/ODM, DPUs/NICs, network operating systems (NOS), optical components, storage platforms, and orchestration software landscape.
- Operational Track Record: Hands-on experience directing production clusters through formal acceptance testing and managing 24/7 mission-critical operations, hardware repairs, and OEM escalations.
- Communication: Exceptional verbal and written communication skills in English, with a demonstrated ability to articulate complex technical trade-offs to internal and external executive stakeholders.
Strongly Preferred
- Experience managing large-scale clusters (500+ nodes or 4,000+ GPUs) and leading teams of 10+ engineers.
- Professional fluency in Business Mandarin to evaluate Chinese-language technical specifications and engage with regional suppliers and clients.
- Direct experience with multi-vendor environments, customer-designed infrastructure, or dual InfiniBand and Ethernet GPU fabric deployments.
First 90 Days Objectives
- 30 Days: Conduct a thorough review of all active APAC designs, BOMs, vendor support agreements, and contractual acceptance obligations; map priority risks and assign technical owners.
- 60 Days: Implement standardized review and validation controls for current project phases; eliminate evidence gaps and align specialized resource planning and hiring strategies.
- 90 Days: Deliver a fully costed regional operating model, publish the initial APAC engineering playbook, and establish a readiness-based vendor management framework.
More Info
Key Skills
GPU-to-NIC topologies
orchestration software
storage platforms
diagnostic frameworks
GPU architectures
NVIDIA GPU-system architectures
