SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Manager, Cluster Engineering & Deployment

TensorWave - Remote - Remote - posted 2026-09-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

TensorWave is building a cloud platform for AI compute at scale. The Senior Manager, Cluster Engineering & Deployment owns the critical function of transforming delivered hardware racks into production-ready clusters. This role is schedule-critical: cluster revenue begins when this team certifies a cluster is ready. Key Responsibilities: - Own and evolve the cluster deployment playbook, including staged bring-up, automated configuration, link/optics validation, cabling verification against port maps, and fault triage during deployment windows. - Lead deployment engineering across concurrent cluster builds through team leads and on-site engineers; coordinate daily with Data Center Integration field teams and cabling vendors. - Drive deployment velocity engineering by reducing cluster bring-up time through tooling, pre-staging, and defect elimination; set team performance targets. - Manage defect feedback loops to Network Engineering, Layer One (cabling), and hardware/optics vendors, holding partners accountable to resolution timelines. - Define spares, test equipment, and deployment tooling requirements per site and standardize across all deployment locations. - Own cluster validation end-to-end: bandwidth/latency baselines, collective (RCCL) performance testing, burn-in criteria, and go/no-go acceptance gates; continuously raise standards as the fleet scales. Requirements: - 10+ years in network deployment, cluster/HPC bring-up, or large-scale infrastructure delivery, including managing engineers in field/deployment settings. - Hands-on fabric bring-up experience at scale (hundreds of switches, thousands of links per deployment). - Strong operational rigor: building and enforcing playbooks, gates, metrics, and blameless defect loops. - Team leadership with schedule accountability across multiple concurrent builds or sites. Preferred Qualifications: - GPU cluster validation experience (NCCL/RCCL benchmarking). - Automation skills (Python, Ansible) applied to deployment. - Optics/link-layer debugging expertise. - Experience with acceptance testing as a commercial gate tied to revenue.

Similar roles