SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Volta is building a vertically integrated AI infrastructure platform with a mission to make compute as dependable and available as electricity. The company has secured $10B in strategic partnerships, Series A funding from Andreessen Horowitz, and a $5B AI Infrastructure Fund, operating across London, Palo Alto, and New York with rapid growth plans.
You will lead Volta's Network Platform Engineering team, responsible for the network layer of the AI infrastructure platform. This is not a traditional ops role—your team writes production software against the fabric, designing and operating large-scale GPU compute infrastructure for AI workloads.
Key responsibilities include:
- Leading a small team of senior network platform engineers: technical direction, design reviews, code review, and delivery
- Setting technical direction for compute and storage fabric, overlay and multi-tenant isolation, edge connectivity, and telemetry/automation
- Owning fabric standards across sites, including RoCE v2 and InfiniBand deployments, maintaining a single reference architecture
- Translating product requirements into scoped technical work and managing roadmap planning
- Designing APIs and abstractions that expose network capability to tenants with isolation and performance
- Collaborating with bring-up teams to surface operational pain points and build scalable platform features
- Evaluating network designs from OEMs and partners, holding vendors to Volta's standards
- Supporting customer-facing technical conversations on workload requirements and fabric design
- Coordinating with security engineering on trust boundaries and tenant isolation
- Owning reliability, observability, and versioning standards for shipped services
- Staying hands-on with complex or high-risk engineering work during bring-up and incident response
- Running post-mortems and closing structural gaps after incidents
- Managing the team through 1:1s, performance conversations, hiring, and onboarding
Required experience:
- 5+ years in large-scale data center or cloud network engineering, with at least 2 years leading an engineering team
- Production software development in Python, Go, or Rust (not scripting)
- Deep Ethernet fabric design and operations: leaf-spine, BGP, ECMP, low-latency production experience
- RoCE v2 at scale: PFC, ECN tuning, DCQCN, understanding GPU training environments
- InfiniBand production experience: fat tree topology, UFM, fabric partitioning, adaptive routing, congestion control
- EVPN and VXLAN in production under multi-tenant load
- Network automation at scale: configuration as code, source of truth systems, CI validation, declarative tooling
- Kubernetes networking: CNI and workload-fabric interactions
- GPU infrastructure experience: understanding training/inference traffic patterns and capacity surfacing
- Multi-vendor credibility across major OEMs
- Security awareness in multi-tenant network design
- Proven people management and clear communication across technical and non-technical stakeholders
- Willingness to be on-site during critical cluster bring-ups
Nice-to-have skills include AI-assisted development, ASN operations, IPv6 at scale, NVIDIA Spectrum-X, gNMI/OpenConfig/NETCONF, OVN/OVS, SR-IOV, DPDK, BlueField DPU, and NVLink/NVSwitch experience.