SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 300,000 / annual
Prime Intellect is building the open superintelligence stack—infrastructure that frontier AI labs build internally, now available to every ambitious AI team. The company's platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment for post-training at frontier scale. Prime Intellect has raised $150M from Founders Fund, Radical Ventures, NVIDIA, and exceptional operators including Andrej Karpathy and leaders from Ramp, Perplexity, Harvey, Cognition, OpenAI, Databricks, and others.
As a Member of Technical Staff - Datacenter Networking, you will design and operate the networks connecting large GPU clusters, owning the reliability and performance of training fabrics, storage networks, and management connectivity so distributed workloads scale without the network becoming a bottleneck.
Core responsibilities include:
- Design and deploy scalable datacenter network topologies for GPU training, inference, storage, and management traffic
- Configure and operate high-performance Ethernet/RoCE and InfiniBand fabrics with clear standards for routing, redundancy, and capacity
- Automate network provisioning, configuration validation, upgrades, and rollback procedures
- Diagnose packet loss, congestion, link failures, and collective communication performance across hosts and switches
- Benchmark end-to-end network performance with infrastructure and ML teams, translating workload needs into measurable acceptance criteria
- Build monitoring for port health, errors, utilization, congestion, and fabric topology; improve incident response and runbooks
- Partner with datacenter operators and hardware vendors on cabling, optics, deployment readiness, and failure resolution
You will work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure, collaborating with a world-class engineering team while having direct impact on systems powering the next generation of AI breakthroughs.
REQUIREMENTS:
Required Experience:
- 3+ years of production datacenter networking experience
- Strong understanding of Ethernet, TCP/IP, routing, switching, and redundant network design
- Hands-on experience with high-performance GPU networking using InfiniBand or RoCE
- Experience troubleshooting network problems across Linux hosts, NICs, switches, and physical links
- Ability to automate network operations with Python, Ansible, or comparable tools
Infrastructure Skills:
- Leaf-spine architectures, BGP, ECMP, VLANs, and network segmentation
- RDMA concepts and performance tuning; congestion control and lossless Ethernet considerations
- Linux networking, NIC drivers and firmware, packet capture, and throughput/latency testing
- Optics, transceivers, cable management, and link-level diagnostics
- Safe change management, configuration versioning, telemetry, and alerting
Nice to Have:
- Experience operating 400G/800G networks or large multi-rack GPU clusters
- NVIDIA Spectrum or Quantum networking experience
- NCCL performance analysis and distributed training troubleshooting
- EVPN/VXLAN, SONiC, or network source-of-truth systems
- Experience with network simulation, automated validation, and capacity planning