Search by job, company or skills

Principal Engineer Cluster Deployment

Early Applicant
  • Posted 13 hours ago
  • Be among the first 10 applicants

Job Description

About Nava

Nava is building next-generation AI infrastructure and inference platforms at global scale. We're looking for a Principal Engineer - Cluster Deployment to lead the successful deployment of large-scale GPU clusters across global data centres.

This is a high-impact execution role responsible for taking GPU clusters from hardware delivery to fully validated, production-ready infrastructure. You'll own deployment execution across sites, vendors, and cross-functional teams, ensuring every cluster is delivered on time, meets quality standards, and is ready for customer workloads.

What You'll Do

  • End-to-End Cluster Deployment
    • Own the deployment of GPU clusters from rack arrival through to tenant-ready production capacity.
    • Plan, coordinate, and execute deployment activities across multiple global data centre locations.
    • Drive installation, rack integration, power-up, hardware bring-up, firmware validation, and production readiness.
    • Ensure deployments are completed safely, efficiently, and within agreed timelines.
  • Deployment Execution
    • Lead cluster bring-up activities, including hardware validation, network configuration, storage integration, and platform initialization.
    • Oversee structured cabling, power validation, firmware upgrades, BIOS configuration, and infrastructure readiness.
    • Coordinate deployment activities across internal engineering teams, hardware vendors, systems integrators, and data centre partners.
    • Resolve deployment blockers and drive rapid issue resolution during implementation.
  • Validation & Quality Assurance
    • Develop and execute deployment validation procedures and acceptance criteria.
    • Lead burn-in testing, stress testing, hardware diagnostics, and production readiness validation.
    • Ensure clusters meet defined performance, stability, and reliability benchmarks before customer handover.
    • Own first-pass acceptance and minimize deployment rework through robust quality processes.
  • Vendor & Site Management
    • Act as the primary technical lead during on-site deployments.
    • Manage relationships with OEMs, contract manufacturers, data centre operators, and installation partners.
    • Ensure deployment standards are consistently followed across all locations.
    • Drive continuous improvement across deployment processes, documentation, and execution methodologies.
  • Cross-Functional Collaboration
    • Partner with GPU Cluster Engineering, Platform Engineering, Networking, SRE, Supply Chain, and Data Centre Operations teams.
    • Ensure seamless transition from deployment to production operations.
    • Support troubleshooting of deployment issues and coordinate engineering fixes where required.
  • Operational Excellence
    • Develop deployment playbooks, SOPs, checklists, and automation to improve deployment consistency.
    • Track deployment metrics and identify opportunities to reduce deployment timelines and improve quality.
    • Drive lessons learned reviews following every deployment.
Success Metrics

You Will Be Measured On

  • Time-to-live for new GPU clusters
  • First-pass deployment acceptance rate
  • Deployment quality and reliability
  • Reduction in deployment rework
  • Deployment schedule adherence
  • Production readiness at handover
  • Deployment process standardization and automation

Qualifications

Required Qualifications

  • 10+ years of experience in infrastructure deployment, data centre engineering, HPC, cloud infrastructure, or systems engineering.
  • Proven experience deploying large-scale compute infrastructure, GPU clusters, HPC systems, or cloud platforms.
  • Deep technical expertise in:
    • Server and rack deployment
    • GPU hardware platforms
    • High-speed networking (InfiniBand, Ethernet, RDMA)
    • Structured cabling and power systems
    • Linux systems administration
    • Firmware, BIOS, and hardware lifecycle management
    • Infrastructure validation and production readiness
  • Experience coordinating cross-functional deployment projects across multiple sites.
  • Strong troubleshooting and problem-solving skills in complex infrastructure environments.
  • Willingness to travel to domestic and international deployment locations.
Preferred Qualifications

  • Experience with NVIDIA DGX, HGX, or similar GPU platforms.
  • Familiarity with Kubernetes, cluster provisioning, automation, and Infrastructure-as-Code.
  • Experience working with hyperscalers, AI infrastructure providers, or data centre operators.
  • Exposure to HPC, AI factories, or large-scale inference platforms.

Skills: drive,platforms,infrastructure,data,gpu,production readiness,automation,cluster,validation,firmware,readiness

More Info

Job Type:
Industry:
Employment Type:

About Company

Job ID: 152254827

Beware of Scammers

We don’t charge money for job offers