

Search by job, company or skills

We are seeking a Compute Cluster / HPC Engineer to design and develop high-performance compute cluster configurations optimized for performance, reliability, and scalability within CLIENTS systems. The role involves selecting, integrating, and validating hardware components such as CPUs, memory, storage, networking, and specialized accelerators, with a strong focus on InfiniBand-based architectures. You will collaborate closely with hardware, software, and systems engineering teams to ensure seamless system integration, participate in design reviews and integration planning, and contribute to cross-functional problem-solving efforts. The position also requires documenting hardware design decisions, integration procedures, diagnostic workflows, and supporting rack-level design considerations including power, cooling, and cabling.
Required Skills & Qualifications
The ideal candidate has solid exposure to HPC environments running SUSE Linux, with hands-on experience in Linux system administration and OS customization. You should be familiar with InfiniBand fundamentals, common IB tools, and troubleshooting practices (candidates are expected to specify the tools they have used), as well as an understanding of CMU and golden image concepts. Experience with rack design fundamentals, FRU replacement or qualification processes, and system-level performance tuning is required, along with a strong understanding of hardware–software interaction. Excellent documentation and communication skills are essential to effectively support internal teams and cross-functional collaboration.
Job ID: 152296403
Skills:
Linux System Administration, Distributed Systems, compute cluster or server environments, hardware validation and troubleshooting tools, system-level performance tuning, computer hardware design, OS customization, hardware-software interaction