Search Jobs

Search by job, company or skills

FIELD APPLICATION ENGINEER, AI FACTORY

FIELD APPLICATION ENGINEER, AI FACTORY

Dimension.ai
5-7 Years
  • Posted 18 hours ago
  • Be among the first 10 applicants

Job Description

ABOUT DIMENSION AI

Dimension AI develops and delivers AI compute infrastructure and related services for enterprise customers. Our work spans GPU infrastructure procurement, data centre deployment, managed services and commercial partnerships across multiple markets. As we build and operate AI Factories for our customers, we seek a hands-on engineer who can turn advanced GPU, networking and software platforms into reliable, high-performing production environments.

THE OPPORTUNITY

The Field Application Engineer, AI Factory, will be Dimension AI's technical lead in the field for AI Factory deployments. The role sits between customers, our service delivery and operations teams, and our technology partners, and covers the full lifecycle from solution design and on-site deployment through performance validation, handover and ongoing support.

The successful candidate will translate customer workloads and business requirements into well-engineered AI Factory architectures and will be trusted by customers and technology partners to solve problems on site, communicate clearly and see deployments through to production.

KEY RESPONSIBILITIES

Customer Engagement and Solution Design

▪ Work with customers and the commercial team to understand AI training, fine-tuning and inference workloads, performance targets, capacity requirements and growth plans.

▪ Design AI Factory reference configurations covering GPU compute, high-speed networking, storage, cluster software, power and cooling, sized to customer requirements.

▪ Prepare technical proposals, architecture documents, bills of materials and responses to customer technical and security questionnaires.

▪ Run technical workshops, proof-of-concept trials and customer demonstrations, and present findings and recommendations clearly to technical and executive audiences.

▪ Advise on data centre readiness, including rack density, power, liquid or air cooling, cabling and network integration.

AI Factory Deployment and Commissioning

▪ Lead on-site installation, configuration and commissioning of GPU servers, InfiniBand or high-speed Ethernet fabrics, storage systems and management networks.

▪ Deploy and configure cluster management, scheduling and orchestration software such as Slurm and Kubernetes, together with GPU drivers, container runtimes and monitoring tools.

▪ Coordinate with data centre, facilities, logistics and OEM teams to keep deployment schedules, site access and equipment deliveries on track.

▪ Maintain deployment runbooks, configuration baselines, as-built documentation and acceptance checklists for each site.

▪ Carry out structured handover to customers and operations teams, including training and documentation.

Performance Validation and Optimisation

▪ Run burn-in, stress and acceptance tests on compute, network and storage, using tools and benchmarks such as NCCL tests, HPL and representative training and inference workloads.

▪ Diagnose and resolve performance bottlenecks across GPU utilisation, interconnect topology, storage throughput, scheduler configuration and software stack.

▪ Establish performance baselines and telemetry so that degradation, thermal issues and hardware faults are detected early.

▪ Help customers move workloads onto the AI Factory and tune frameworks, containers and job configurations for efficient use of the cluster.

Operations, Support and Escalation

▪ Act as the senior technical escalation point for customer incidents, performing root cause analysis and driving issues to resolution.

▪ Work with NVIDIA, OEM and network and storage vendors on fault isolation, firmware and software updates, and RMA cases.

▪ Contribute to operations and maintenance SOPs, troubleshooting guides, change procedures and incident reports.

▪ Support the operations team with maintenance windows, upgrades and capacity expansions, and record lessons learned after each event.

▪ Work within export control, customs and data centre access requirements that apply to GPU systems and related equipment.

Technical Partnerships and Continuous Improvement

▪ Build strong working relationships with technology partner engineers and OEM field teams, and stay current on new GPU platforms, networking, cooling and AI software developments.

▪ Provide field feedback to the commercial, procurement and product teams on customer requirements, deployment risks and design improvements.

▪ Develop standard designs, automation scripts and deployment tooling that make each new AI Factory faster and more repeatable to deliver.

▪ Share knowledge with colleagues and mentor junior engineers and technicians as the team grows.

CANDIDATE PROFILE

▪ Bachelor's degree in Computer Engineering, Electrical Engineering, Computer Science or a related discipline, or equivalent practical experience.

▪ Significant experience in a field application, solutions, systems or infrastructure engineering role involving GPU or high-performance computing environments.

▪ Hands-on experience deploying and troubleshooting GPU servers and clusters, including drivers, CUDA-based software stacks, containers and cluster management tools.

▪ Working knowledge of high-speed networking such as InfiniBand, RoCE or Ethernet fabrics, and of high-performance storage used in AI and HPC environments.

▪ Proficiency in Linux administration and scripting (for example, Bash or Python), and familiarity with Slurm, Kubernetes and infrastructure automation tools.

▪ Understanding of data centre fundamentals, including power distribution, cooling and rack-level design; exposure to liquid cooling is an advantage.

▪ Clear written and verbal communication, with the confidence to work with customers, executives, vendors and on-site contractors.

▪ A practical, self-directed approach, sound judgement under pressure, and willingness to travel and work on site, including across the Asia-Pacific region.

Relevant certifications from technology partners or in Linux, networking or cloud are advantageous, as is experience with large-scale AI training or inference clusters, MLOps platforms or managed services environments. Awareness of export control and cross-border equipment movement requirements is also valuable.

WHAT SUCCESS LOOKS LIKE

Within the first year, the successful candidate will have led one or more AI Factory deployments from design through customer handover, established repeatable deployment and acceptance standards, and built strong working relationships with customers and technology partners. Customers will have production-ready environments that perform to specification, and Dimension AI's operations team will have clear documentation, baselines and escalation paths to support them.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

scheduling and orchestration software

high-speed Ethernet fabrics

GPU drivers

container runtimes

Slurm

GPU servers

About Company

Similar Jobs

3-5 yrs
SGD 4,500 - 6,000 per month
Singapore
Skills:
ServersNetworkingStorageData Center Infrastructureapplications engineeringpower solutionsconnector technologiesinterconnect systems
8-10 yrs
Singapore
Skills:
cloud hosting Kubernetessoftware librariesSDKscontainer runtimesnimshypervisorsAI Factory infrastructureproduct integrationsNVIDIA software stack
5-7 yrs
Singapore
Skills:
OpenmpLinux AdministrationCudaHPC application performance testingTechnical program managementOpenACCPerformance debug testingTCO modelsInstinct GPUsProof of ConceptsAMD EPYC CPUs
3-5 yrs
Singapore
Skills:
PythonKubernetesIscsiNasSannetwork storageSRE principlesUDP protocol stack
5-7 yrs
Singapore
Skills:
CGithubJiraHubspotPythonConfluenceOscilloscopeslogic analyzersSensorshardware debugging toolsedge AI platformsCommunication ProtocolsEmbedded Systems