Search by job, company or skills

AI Infrastructure Support Engineer

3-5 Years
SGD 7,500 - 12,500 per month
  • Posted 7 hours ago
  • Be among the first 10 applicants

Job Description

What You'll be Doing

  • Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately. Collaborate with Engineering, with guidance, when incidents or changes require it
  • Perform GPU node triage and hardware troubleshooting: interpret nvidia-smi/DCGM output and system logs, isolate faults across GPU, NIC, and server hardware, carry out physical remediation (reseats, swap testing, component checks), and prepare clean evidence for vendor RMA
  • Run fabric and link diagnostics following established runbooks (mlxlink orequivalent), capture evidence accurately, and escalate with a handover that lets the next engineer continue without starting from scratch
  • Assist with storage and data-path investigations (mounts, connectivity, client-side symptoms) on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis
  • Follow established runbooks to resolve common issues propose improvements and contribute incremental fixes with review
  • Accurately record, update, manage, and resolve tickets, keeping all parties informed with clear notes, next steps, and customer communications via the agreed channels
  • Participate in monitoring, troubleshooting, and triage. Capture logs and facts to enable efficient handover
  • Participate in changes under peer review, learning risk assessment and backout practices in live customer environments
  • Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns)
  • Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes
  • Be the escalation point for onsite DC Operations staff coordinate smart- hands tasks within your scope
  • Learn the Platform fundamentals so you can help customers get value from our services, asking for support when deeper expertise is needed
  • Share knowledge by documenting steps you've validated and contributing to training materials.
  • Take part in incident reviews as a contributor and help track preventative follow-ups in your scope
  • Deliver assigned tasks and project work to agreed quality and timelines. Flag blockers early and seek help when needed
  • Participate in on-call and out-of-hours work when scheduled and after onboarding. Travel to Nscale or customer locations to assist with deployments, troubleshooting, and operational tasks, and attend supplier training as required.

About You

  • Experience. 3+ years in infrastructure support or support engineering, including support/service desk experience in structured, SLA-driven, customer-facing environments (cloud, data centre, or managed services)
  • Communication. Clear written notes, concise updates, and reliable follow- through. Able to explain technical issues accurately to customers and colleagues, and produce handovers the next shift can act on immediately
  • GPU and hardware troubleshooting. Working knowledge of GPU infrastructure: hands-on with nvidia-smi or similar diagnostics, comfortable interpreting hardware error output and logs, and confident physically troubleshooting servers - reseating components, swap testing, working via
  • BMC/out-of-band management - through to preparing RMA evidence. A strong technical base here is required, not a learning goal
  • Linux. Solid working knowledge: confident on the CLI with systemd, filesystems, permissions, and standard networking tools. Able to troubleshoot common issues independently and know when to escalate
  • Networking. Solid grasp of IP addressing, subnets, VLANs, routing, DNS, and firewalls. Awareness of high-performance east-west fabrics (RDMA/InfiniBand concepts) is a plus and a core growth area in this role
  • Ticketing and ITSM discipline. Experience working within structured support processes (ITIL or similar): prioritisation, escalation, SLA awareness, and accurate documentation
  • Observability foundations. Able to use dashboards and alerts to identify symptoms, gather evidence, and follow runbooks. Comfortable proposing simple alert or dashboard improvements with review
  • Scripting and automation basics. Comfortable reading and writing simple Bash or Python, and using Git for version control
  • Platform and DC fundamentals. Understanding of servers, networks, storage, and virtualisation concepts, ideally from a support or operations background
  • Growth mindset. Curious, dependable, and collaborative. You seek feedback, ask questions, and invest in learning to progress toward Senior
  • Adaptability. Able to work in a fast-moving environment with evolving processes, participate in on-call after onboarding, and travel when needed

More Info

Job Type:
Industry:
Employment Type:

Job ID: 153609863

Similar Jobs

Singapore, Marina

Skills:

PrometheusBashGrafanaBGPFirewallsVlansTerraformAnsibleLoad BalancingVxlanPythonnvidia-smiinfinibandMAASRoCEXID error interpretationBMC RedfishLinux systems engineeringHPC schedulingibdiagnetmlxlinkDCGMSlurm

Beware of Scammers

We don’t charge money for job offers