Search Jobs

Search by job, company or skills

3-5 Years
SGD 5,000 - 7,500 per month
Early Applicant
  • Posted a month ago
  • Be among the first 10 applicants

Job Description

We are seeking an experienced AI System Engineer to design, deploy, operate, and optimize AI training clusters, GPU computing platforms, and supporting infrastructure. The ideal candidate should possess strong expertise in Linux systems, GPU computing environments, container platforms, and AI/HPC cluster architectures.

Key Responsibilities

AI Cluster deployment and operation

. Deploy and operate AI training and HPC clusters

. Install, configure, and optimize operating systems on GPU servers

. Manage cluster resources and capacity

. Perform system upgrades, patch management, and change implementation

. Develop and maintain standardized operational procedures.

Linux System Management

. Manage large-scale Linux environments

. Perform system performance tuning

. Analyze system logs and kernel issues

. Troubleshoot system stability problems

. Manage user access and security policies.

GPU Platform Support

. Manage NVIDIA GPU computing platforms

. Deploy and maintain CUDA, NVIDIA Drivers, and Fabric Manager

. Troubleshoot GPU, NV Link, and NV Switch-related issues

. Optimize GPU cluster performance

. Support customer in resolving training environment issues.

Container Platform & Orchestration System

. Build and maintain Kubernetes clusters

. Support AI workload scheduling

. Deploy and manage container runtime environments

. Optimize GPU utilization within containers

. Manage Kubernetes high-availability architectures.

AI infrastructure management

. Manage distributed storage platforms

. Operate high-speed networking environments

. Collaborate with datacenter teams for troubleshooting

. Monitor infrastructure health and performance

. Improve system reliability and availability.

Automated operation and maintenance and platform development

. Develop infrastructure automation tools

. Create deployment and health-check scripts

. Build monitoring and observability platforms

. Implement alerting and self-healing mechanisms

. Improve operational efficiency through automation.

  • Good communication, teamwork, and ownership mindset.
  • Willing to participate in on-call rotation, maintenance windows, and emergency incident response, willing to accept short-term business trips.

Required Qualifications

  • Bachelor's degree or above in Computer Engineering, Electrical Engineering, Telecommunications, or related fields.
  • Linux System
    . 3+years of Linux administration experience
    . Strong knowledge of Ubuntu, Rocky Linux, and RHEL . Familiarity with system boot process, kernel, filesystems, and performance tuning
    . Ability to troubleshoot complex system issues independently.
  • GPU & AI Platform

    Strong understanding of NVIDIA GPU architecture
    Experience with NVIDIA GPU products H100, H200, B200 B300 ,GB200 NVL72 and GB300 NVL72
  • Familiar with CUDA, NCCL, NV Link, NV Switch, GPU Direct RDMA
  • Understanding of distributed AI training architectures.

Container and Cloud Native

. Hands-on experience with Kubernetes

. Familiarity with Docker and Containerd

. Experience with Helm

. Knowledge of GPU Operator

. Understanding of Kubernetes GPU scheduling.

Networking and Storage

. Strong understanding of TCP/IP networking

. Experience with InfiniBand and RoCE

. Familiarity with RDMA architectures

. Experience with one or more storage systems Lustre, BeeGFS, Ceph

and NFS

Automation capabilities

. Strong scripting skills in Shell and Python

. Know about Ansible

. Ability to build infrastructure automation scripts.

Preferred Qualities


. Experience supporting Large Language Model(LLM) training platforms

. Knowledge of Slurm workload manager

. Experience with Ray and Kubeflow

. Familiarity with NVIDIA Base Command Manager(BCM)

. Experience with NVIDIA NIM

. Knowledge of PXE deployment solutions

. Experience on using DDN product

. Experience operating large-scale GPUclusters

. Experience supporting global datacenteroperations.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

RDMA

NVIDIA Drivers

TCP IP networking

RoCE

NV Switch

container platforms

GPU Direct

BeeGFS

Fabric Manager

GPU computing environments

AI HPC cluster architectures

NVIDIA GPU architecture

Linux systems

NCCL

NV Link

Similar Jobs

2-4 yrs
Singapore
Skills:
VMware, security practices, Windows Server, Storage Systems, Monitoring Tools, Network Protocols, Linux, Hyper-V, troubleshooting methodologies, virtualization technologies, firewall configurations, server installation and management
2-5 yrs
SGD 6,000 - 7,000 per month
Singapore, Ang Mo Kio
Skills:
Group Policy Objects (GPO), Linux-based Virtual Appliances, Network and network security solutions, DevOps operations, Microsoft SQL Server 2019, Commvault backup jobs, Windows RDS Architecture, Ubuntu OS, Kubernetes platform, Microsoft Failover Clustering, Windows Domain Administration, Windows Server 2019, SSL certificate management, AzureStack Hub Virtualisation Environment, Ansible automation tools
2-5 yrs
SGD 4,000 - 5,500 per month
Singapore, Ang Mo Kio
Skills:
Spring Boot, Node.js, Linux Administration, Devops, React, Vulnerability Management, Docker, Kubernetes, middleware platform support, CI CD pipeline management, NoSQL databases, deployment automation, Certificate and secret management, Secure system hardening, production support operations, Troubleshooting, Podman, SQL databases, Cybersecurity best practices, container technologies
5-8 yrs
SGD 5,000 - 7,000 per month
Singapore
Skills:
Unix, System Integration Testing, test automation, Test Execution, Confluence, Defect Management, Performance Testing, Test Design, Test Planning, Performance Tuning, Jira, User Acceptance Testing, Linux, microservices security validation, Windows operating systems, API security testing automation, Automatic Fare Collection systems, DevSecOps frameworks, automated testing pipelines, cloud-based testing environments, scripting or programming languages, Micro-payments, Embedded Systems, Xray, GitLab CI CD
2-5 yrs
SGD 4,000 - 5,500 per month
Singapore, Jurong East
Skills:
telecommunications systems , Visio), Microsoft Office (Word, can bus , Excel, System integration, Vehicle electrical electronic architecture and integration, Optical test tools, EMI EMC considerations, Oscilloscopes, TCP/IP networking, Spectrum analyzers, Vehicle electronics, Electro-optical systems, RF communications systems, Embedded Systems, EO IR cameras, Embedded systems control, Powerpoint, Sensors, signal generators, environmental testing, Test and measurement equipment