Our client is a fast-growing AI infrastructure company building next-generation GPU-powered AI platforms. They operate high-performance AI data centers that support large-scale machine learning, AI model training and inference workloads. The environment is highly technical, focusing on low-latency networking, scalability, automation and operational excellence.
Primary Responsibilities
- Design, deploy, and support high-performance data center network infrastructure for AI and GPU compute environments.
- Build, configure, and optimize Spine-Leaf network architectures for large-scale GPU clusters.
- Deploy and manage high-speed Ethernet and/or InfiniBand fabrics supporting AI/HPC workloads.
- Configure and maintain Layer 2 and Layer 3 networking technologies, including BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC.
- Monitor, troubleshoot, and resolve complex network issues across compute, storage, and AI infrastructure.
- Optimize network performance, latency, and throughput to support distributed AI training and inference.
- Perform firmware upgrades, network maintenance, capacity planning, and lifecycle management.
- Collaborate closely with Platform, Infrastructure, Linux, DevOps, and AI Engineering teams to deliver scalable AI infrastructure.
- Develop and maintain network documentation, operational procedures, and technical runbooks.
- Implement network monitoring, alerting, and automation to improve operational efficiency.
- Participate in incident response and on-call support for production environments.
- Evaluate and recommend new networking technologies to improve scalability, reliability, and performance.
What We're Looking For
- Bachelor's Degree in Computer Science, Computer Engineering, Information Technology, or a related discipline.
- 5+ years of experience designing or supporting enterprise or data center network infrastructure.
- Strong understanding of modern data center networking principles
- Hands-on experience with routing and switching protocols such as BGP, OSPF, VXLAN EVPN, ECMP, and MLAG/VPC.
- Experience managing high-speed Ethernet networks (25G/40G/100G/200G/400G).
- Exposure to AI, HPC, GPU clusters, or large-scale compute environments would be highly advantageous.
- Experience working with networking platforms such as Cisco Nexus, Arista, Juniper, NVIDIA Spectrum, or Mellanox.
- Good understanding of network security concepts including segmentation, ACLs, and firewall policies.
- Familiarity with Linux networking fundamentals and troubleshooting.
- Experience with network automation using Python, Ansible, REST APIs, or similar tools is an advantage.
- Knowledge of technologies such as InfiniBand, RoCEv2, RDMA, GPUDirect, SONiC, or Cumulus Linux is a plus.
- Strong analytical, troubleshooting, and problem-solving skills.
- Excellent communication skills with the ability to collaborate effectively across cross-functional engineering teams.
- Comfortable working in a fast-paced, high-availability production environment supporting mission-critical AI infrastructure.
Click on Apply now to find out more about this opportunity and other available positions.
EA License: 22C1396
EA Personnel: R1551466