SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
xAI is seeking an experienced Network Engineer to design, build, and operate mission-critical networks powering AI supercomputer campuses. You will work on a small, highly motivated infrastructure team focused on engineering excellence in a flat organizational structure where all employees contribute directly to the company's mission.
Key responsibilities include designing and implementing highly available, low-latency, high-bandwidth networks for AI training and inference clusters; evaluating and deploying data-center class network hardware (switches, NICs, firewalls, optical multiplexers) supporting 400G/800G speeds; contributing to network automation tooling using GitOps and Infrastructure as Code frameworks; and troubleshooting network issues affecting cluster performance.
You will coordinate network change windows with stakeholders, provide direct support during cluster bring-up and expansion, implement proactive monitoring and telemetry for fabric health and congestion detection, and maintain comprehensive network documentation. The role requires collaboration with infrastructure, compute, storage, site operations, and enterprise teams to identify and resolve design issues, especially failure modes in AI fabrics.
Additional responsibilities include ensuring compliance with industry and cybersecurity standards (ITAR, ISO, NIST), performing site walks with customers and vendors, and serving as an on-call networking engineer during operations. Work may include evenings and weekends aligned with compute schedules.
Required qualifications: Bachelor's degree in computer science, computer engineering, or STEM discipline plus 3+ years of professional network engineering experience (or 5+ years without degree); extensive hands-on experience with Layer 2/3 networks in latency-sensitive and data-center environments; functional experience with multiple network vendors; and experience with GitOps and Infrastructure as Code frameworks.
Preferred experience includes strong OSI model knowledge, hands-on work with Cisco, Arista, Juniper, or NVIDIA Spectrum-X switches, RoCEv2 Ethernet AI/HPC fabric experience, and familiarity with AI training/inference traffic patterns and NCCL.