SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 250,000 / annual
xAI is building large-scale networks that underpin training and inference infrastructure for AI systems. This role is a hands-on network engineering position focused on production network design, deployment, and operations at scale.
You will own the full lifecycle of datacenter and campus/core networks, including:
- Designing and deploying production networks at scale
- Owning routing and switching configuration standards (BGP, OSPF, IS-IS), including change design, peer reviews, and execution
- Qualifying new network platforms, optics, and topologies; contributing to architecture and capacity planning
- Building and improving monitoring, alerting, and operational documentation
- Troubleshooting Layer 2/Layer 3 incidents end-to-end (link flaps, optics, routing, traffic engineering) and driving root cause analysis
- Automating repetitive network tasks using Python, Ansible, or similar tools
- Partnering with compute, facilities, and software teams during cluster build-outs and maintenance windows
- Supporting high-performance/supercompute network environments (Ethernet AI/HPC fabrics, RoCE/RDMA-capable designs)
Travel to Memphis and other build sites may be required for capacity build-outs. You will participate in a team on-call rotation. The organization operates with a flat structure; all employees are expected to be hands-on and contribute directly to the mission.
REQUIREMENTS:
- Several years designing and/or operating production networks in a datacenter, ISP, cloud, or large enterprise environment
- Solid hands-on experience with BGP and at least one interior routing protocol
- Working knowledge of TCP/IP, VLANs, EVPN/VXLAN or equivalent datacenter overlays, and optics/high-speed Ethernet
- Experience troubleshooting live production network incidents and participating in on-call
- Strong written and verbal communication skills; ability to write clear change docs and incident notes
- Willing to work onsite in Palo Alto
PREFERRED SKILLS:
- Experience with modern datacenter vendors (Arista, Cisco, Juniper, Nvidia/Mellanox)
- Familiarity with high-performance or supercompute networking (RoCEv2, congestion control, GPU cluster fabrics)
- Network automation (Python, Ansible, Terraform, or similar) used in production
- Experience with EVPN, leaf-spine, and large-scale Ethernet fabrics
- Prior work supporting rapid datacenter or cluster capacity build-outs