Search Jobs

Search by job, company or skills

Senior Distributed Systems & AI Infrastructure Engineer

Senior Distributed Systems & AI Infrastructure Engineer

nexadept
Fresher
  • Posted 17 hours ago
  • Be among the first 10 applicants

Job Description

About the Role

We are looking for a Senior Distributed Systems & AI Infrastructure Engineer to build and optimize the infrastructure powering next-generation AI inference at global scale.

You will work on high-performance distributed systems supporting large-scale model serving across thousands of GPUs and other accelerators. The role sits at the intersection of distributed systems, performance engineering, networking, and AI infrastructure, with a strong focus on building systems that are fast, scalable, and highly reliable.

Key Responsibilities

  • Design and develop high-performance distributed systems for large-scale AI inference.
  • Build scalable infrastructure for model serving across large GPU and accelerator clusters.
  • Optimize system latency, throughput, resource utilization, and reliability.
  • Design and improve high-performance networking and I/O components.
  • Investigate and resolve complex issues across distributed systems and large-scale infrastructure.
  • Develop performance-critical components using Rust, Go, or C++.
  • Improve the scalability and reliability of infrastructure supporting next-generation AI workloads.
  • Work closely with infrastructure, ML, and systems teams to identify and solve performance bottlenecks.

Requirements

  • Bachelor's degree or equivalent experience in Computer Science, Engineering, or a related technical field.
  • Strong systems programming experience in Rust, Go, or C++.
  • Proven experience designing and building high-performance distributed systems at scale.
  • Strong understanding of networking, network protocols, and high-performance I/O.
  • Strong debugging and problem-solving skills for complex distributed systems.
  • Experience optimizing systems for performance, scalability, and reliability.

Preferred Qualifications

  • Experience with AI/ML serving infrastructure or large-scale inference systems.
  • Familiarity with disaggregated inference architectures.
  • Understanding of GPU programming models and GPU memory hierarchies.
  • Experience with GPU networking and interconnect technologies such as NVLink, InfiniBand, or RoCE.
  • Experience working with large-scale GPU clusters or accelerator infrastructure.
  • Knowledge of performance optimization, profiling, and systems benchmarking.
  • Experience supporting large-scale AI model training or inference workloads.

More Info

Job Type:
Industry:
Employment Type:

Key Skills

GPU memory hierarchies

RoCE

GPU programming models

GPU networking

high-performance distributed systems

NVLink

systems benchmarking

About Company