SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Volta is building a vertically integrated AI infrastructure platform with a mission to make compute as reliable and available as electricity. The company operates large-scale GPU compute clusters—each containing tens of thousands of GPUs, hundreds of switches, and tens of thousands of cables—requiring sophisticated automation and modeling to manage effectively.
In this role, you will own the network source of truth: a comprehensive data model covering sites, racks, devices, interfaces, cabling, addressing, and topology. You'll generate device configuration from this model across all platforms in the estate, ensuring configuration is always an output of the model rather than manually edited. You will build and operate a simulation environment that reproduces full cluster topologies, enabling pre-deployment validation of fabric changes as a standard practice.
Key responsibilities include building CI pipelines that validate network changes through schema checks, policy validation, generated config diffs, simulated convergence, and reachability assertions before merge. You'll detect and close drift between intended and actual device state across sites, making divergence visible rather than discovered during incidents. You'll automate bring-up verification with cable plan generation, LLDP-based validation, link quality checks, and acceptance test suites. You'll also build tooling to transform site designs into provisioned fabrics, instrument the fabric with streaming telemetry and topology-aware metrics, and support incident response using modeling and configuration history.
You'll write production Python or Go in shared repositories under the same review, testing, and CI standards as the rest of platform engineering. This is a hands-on role requiring deep network fundamentals and the ability to operate at scale where anything not automated does not happen.