Search by job, company or skills

Senior AI Infrastructure Support Engineer

6-8 Years
SGD 7,500 - 12,500 per month
  • Posted 8 hours ago
  • Be among the first 10 applicants

Job Description

What You'll be Doing

  • Join the Support duty rotation as a senior escalation point, collaborating with Infrastructure Engineering, CNPRE, Network Operations, and Product Engineering on incidents, investigations, and changes
  • Diagnose and remediate GPU node faults across the full stack - driver, firmware, and hardware layers - from nvidia-smi/DCGM and XID/RAS analysis through BMC/Redfish and out-of-band management to physical fault isolation and vendor RMA
  • Own east-west fabric health: run link-level diagnostics (mlxlink, ibdiagnet, or equivalent), isolate transceiver, optics, cabling, and switch-port faults, and validate topology across InfiniBand and RoCE/high-speed Ethernet fabrics.
  • Investigate data-path issues on high-performance storage platforms (e.g. VAST), including storage-network interactions across clients, mounts, VIPs, and routing
  • Run structured, hypothesis-driven investigations conduct root cause analysis for major incidents and drive long-term fixes to completion
  • Author and execute changes in live customer environments with proper risk assessment, peer review, and backout plans
  • Proactively improve dashboards, alerts, and runbooks to prevent repeat incidents identify recurring patterns and convert them into problem records and automation
  • Accurately record, update, and resolve tickets, keeping internal and external parties informed with clear customer-impact statements and evidence-rich notes that enable clean handover
  • Design and implement automation scripts and small tools to reduce toil and human intervention
  • Act as a key escalation point for the Support Organisation taking ownership of strategic decisions where results matter
  • Mentor and upskill mid-level engineers contribute to knowledge sharing across Operations and Engineering, including training content, workshops, and PR reviews
  • Lead by earning trust and speaking candidly. Disagree when appropriate and challenge the status quo commit wholly to decisions once in motion
  • Respond to critical incidents out of business hours and participate in on-call as required. Travel to Nscale or customer sites to provide onsite technical expertise

About You

  • Experience. 6+ years in infrastructure, operations, or support engineering in production environments 2-3+ years hands-on with GPU, HPC, or large- scale data centre estates, ideally in a customer-facing or escalation-driven capacity
  • Communication. Able to explain complex technical detail clearly, specifically, and concisely - in tickets, in incident updates, and face to face with customers and stakeholders at all levels. Strong written discipline: your notes let the next engineer pick up where you left off without starting from scratch
  • GPU platforms (NVIDIA AMD Instinct beneficial). Practical, current experience with GPU drivers, firmware, and runtime stacks on AI training and inference clusters. Confident with nvidia-smi, DCGM, and XID/error interpretation able to isolate faults across GPU, baseboard, NIC, and PCIe layers and drive them through diagnosis to RMA
  • High-performance east-west fabrics. Hands-on experience with RDMA fabrics such as InfiniBand and/or RoCE including link-layer diagnostics (mlxlink, ibdiagnet, or equivalent), transceiver and cabling fault isolation, and understanding of rail-optimised topologies, NVLink/NVSwitch, and NCCL- based performance troubleshooting on multi-node clusters
  • HPC scheduling. Slurm operations for large multi-GPU jobs - containers via Pyxis/Enroot, MPI, and diagnosing queue, topology, and job failures
  • Linux systems engineering at scale. Strong command of modern Linux distributions, kernel modules, systemd, networking stack, and filesystem tooling. Proven troubleshooting across compute, storage, and network layers in production
  • Server hardware and control planes. Comfortable with BMC/Redfish, firmware management, and bare-metal provisioning workflows (MAAS or similar) across large node fleets
  • Networking fundamentals. Solid grasp of L2/L3, routing, BGP, VLANs, VXLAN, firewalls, and load balancing, with a clear understanding of how east- west cluster traffic differs from north-south
  • Observability and incident response. Build and use alerting stacks and dashboards (Prometheus/Grafana or similar), interpret metrics and alerts, drive runbooks to resolution, and contribute to SLOs and post-incident reviews
  • Change and risk judgment. Experience authoring and executing changes in business-critical environments, including risk assessments, customer-impact analysis, and backout plans
  • SRE-style operations. Write and maintain runbooks, automate diagnostics, and reduce human intervention through scripts and small tools
  • Automation and Git. Scripting skills in Bash, Python, or equivalent for operational tooling and integrations experience with infrastructure automation tools (Ansible, Terraform, or similar)
  • Data Centre fundamentals. Understanding of how data centres operate - servers, networks, storage, power, and cooling ideally gained through an operational support background
  • Leadership. Disciplined, organised, and self-motivated, with the ability to mentor and motivate other engineers, take decisive action, and drive the team and wider organisation to improve
  • Adaptability. Able to adapt to customer-driven demands, including specialist support outside core hours and travel for onsite work

More Info

Job Type:
Industry:
Employment Type:

Job ID: 153609891

Beware of Scammers

We don’t charge money for job offers