About Nava
Nava is building Asia's next-generation AI-native cloud platform—designed to empower the world's most ambitious AI developers with scalable, high-performance infrastructure. Our mission is to democratize access to enterprise-grade AI compute by delivering a secure, efficient, and developer-friendly cloud experience. While our technology is deeply rooted in systems innovation, we believe that great infrastructure should be invisible: enabling builders to focus on what matters most—their models, not the hardware.
Role Responsibilities
We are rapidly scaling our AI-native GPU cloud to support massive-scale inference and training workloads across strategic Tier-1 APAC hubs. As we deploy multi-megawatt, high-density AI infrastructure, we're looking for a
Head of Cluster Engineering to lead the design, orchestration, driver/firmware integration, and scaling of our multi-tenant accelerator fabrics.
This is a critical leadership role where you'll bridge deep technical expertise with strategic vision—ensuring our infrastructure delivers the performance, reliability, and flexibility needed to power next-generation AI applications. You'll own the core systems that power our low-latency inference engine, high-performance training platforms, and the Nava Data Platform.
What You Will Do
- Engineering Leadership: Build, scale, and mentor an elite engineering team dedicated to control plane architecture, deep systems tuning, driver/firmware automation, and cluster-scale orchestration.
- Firmware & Driver Lifecycle: Direct automated bare-metal provisioning, BIOS/UEFI configurations, BMC/IPMI integrations, firmware flashing pipelines, and host GPU driver stack deployment across heterogeneous accelerator hardware.
- NVIDIA Stack & Accelerator Integration: Drive deep hardware-software co-design across the NVIDIA ecosystem (CUDA, NVLink/NVSwitch, NCCL, GPUDirect Storage/RDMA, BlueField DPUs, DOCA, TensorRT) to maximize collective communication bandwidth and GPU utilization.
- Advanced Orchestration: Drive the development and tuning of multi-tenant workload scheduling using Kubernetes, Slurm, and Volcano schedulers for topology-aware, fabric-conscious resource allocation.
- Network & Interconnect Optimization: Optimize high-performance fabrics, leveraging InfiniBand, RoCE v2, RDMA, eBPF, and Cilium to eradicate network latency and tail congestion in large-scale distributed AI workloads.
- Data Platform Interoperability: Seamlessly integrate infrastructure with the proprietary Nava Data Platform, tuning parallel file systems and distributed data pipelines for massive-scale training and low-latency inference.
- Operational Excellence & Observability: Establish declarative GitOps infrastructure pipelines, robust SRE practices, and high-cardinality telemetry spanning hardware health, firmware status, thermals, interconnect error counters, and cluster performance metrics.
What We Are Looking For
- Technical & Systems Mastery: Deep, hands-on architectural experience with large-scale bare-metal provisioning, Linux kernel internals, GPU host drivers, and low-level firmware management.
- Ecosystem Expertise: Expert-level command of NVIDIA system architectures (NVLink topologies, NVSwitch fabrics, CUDA driver/runtime, NCCL tuning, DOCA/DPU offloading) alongside eBPF/Cilium networking and parallel storage systems.
- Leadership Experience: Proven track record of building, managing, and scaling high-performing systems, kernel, network, or cluster engineering teams in production-critical environments.
- Complex Problem Solving: Demonstrated ability to diagnose and solve complex cross-stack failures across driver/firmware boundary conditions, PCIe topologies, network transport layers, and distributed scheduler queues.
- Execution Focus: A relentless drive for delivering highly available, deterministic, and scalable infrastructure capable of reliably backing bleeding-edge AI workloads.
Ideal Background
While experience at hyperscalers, GPU cloud providers, HPC organizations, or AI infrastructure startups is valuable, we welcome strong candidates from diverse technical backgrounds—including enterprise infrastructure, systems software, or even adjacent domains like fintech or robotics—if you can demonstrate deep systems thinking and a passion for AI infrastructure.
Qualifications
Required
- 10+ years of experience in systems, infrastructure, or cluster engineering, with at least 5 years in a senior leadership role.
- Deep expertise in Linux kernel internals, GPU driver stack (NVIDIA CUDA, UVM, persistence mode), and low-level firmware management (UEFI, BMC, IPMI).
- Extensive experience with NVIDIA ecosystem tools and frameworks: CUDA, NCCL, NVLink/NVSwitch, DOCA, GPUDirect, TensorRT, and BlueField DPUs.
- Proven success designing and scaling high-performance, multi-tenant GPU clusters for AI training/inference workloads at scale.
- Strong background in orchestration systems: Kubernetes (K8s), Slurm, or Volcano schedulers—especially with topology-aware scheduling and fabric-aware resource allocation.
- Hands-on experience optimizing high-speed interconnects: InfiniBand, RoCE v2, RDMA, and modern networking stacks (eBPF, Cilium).
- Experience with GitOps, infrastructure-as-code (Terraform, Ansible), and observability tooling for hardware-level telemetry (e.g., health monitoring, error counters, thermals).
Preferred (Nice-to-Have)
- Experience building or operating large-scale AI cloud platforms or GPU infrastructure at hyperscale providers (e.g., AWS, GCP, Azure, CoreWeave, Lambda Labs).
- Familiarity with parallel file systems (e.g., Lustre, BeeGFS, WekaIO) and distributed data pipeline optimization for AI training.
- Contributions to open-source projects related to kernel networking, GPU drivers, or cluster orchestration.
- Background in HPC, high-frequency trading infrastructure, or cutting-edge AI research engineering.
The Opportunity: This is a rare opportunity to lead the foundational cluster engineering strategy for a rapidly expanding AI cloud platform. You will have immense autonomy to shape our technological trajectory, build out an elite engineering team, and define next-generation GPU cloud infrastructure.
Why Join Us
At Nava, you'll be part of a mission-driven team building infrastructure for the next wave of AI innovation backed by top-tier investors and with deep roots in both systems engineering and product thinking. We value curiosity, ownership, and collaboration and we're committed to creating an environment where engineers thrive, grow, and ship impactful technology.
Skills: training,kernel,nvidia,infrastructure,data,cloud,orchestration,cluster,cuda,firmware