SlipstreamJobsFresh Startup & VC-Backed Jobs

Engineering Manager, Deployment

Crusoe - San Francisco, CA, United States - In-office - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 200,000 - 240,000 / annual

Crusoe is a vertically integrated AI infrastructure company building the physical and logical foundation for global AI compute at hyperscale. The company owns and operates every layer of the stack—from energy generation to cloud services—to power the world's most ambitious AI workloads. As Engineering Manager, Deployment, you will lead a team responsible for physically and logically bringing Crusoe's global network of high-performance compute (HPC) and GPU-based AI infrastructure online. You'll manage a deployment team's people, processes, and delivery across a portion of the growing portfolio of concurrent global builds, working closely with Network Deployment technical leadership to execute strategy. Key responsibilities include: - Build and grow your team: Hire, develop, and manage Network Production Engineers across multiple sites and time zones. - Drive deployment automation adoption: Execute automation strategy across the deployment lifecycle, including ZTP-based provisioning, Ansible-driven configuration, and CI/CD pipelines for network turn-up. Hold engineers accountable for replacing manual steps with automated workflows. - Deliver on-time, zero-defect site build-outs: Be accountable for your assigned portion of the portfolio, managing operating cadence and staffing within the broader program. - Execute against strategy: Work with technical leadership to translate deployment strategy and standards into execution plans, and surface field realities that inform standard evolution. - Provide program reporting and risk escalation: Report status to Data Center Ops, Supply Chain, and Site Reliability leadership. - Manage vendor and partner performance: Hold structured cabling vendors, remote hands, and data center providers to defined SLAs and quality standards. Escalate systemic issues. - Champion validation rigor: Ensure burn-in and site acceptance testing (SAT) standards are consistently met. Partner on handoff quality and root-cause resolution. - Support capacity planning: Provide input on hardware lifecycle, lead-time risk, and turn-up cadence. - Contribute to continuous improvement: Participate in cross-site retrospectives and post-incident reviews. Surface gaps in tooling, staffing, or process. - Develop team members: Coach engineers into stronger technical and people contributors. - Apply network reference architectures: Build sites against established architectures for frontend GPU/compute fabric (leaf-spine design, RDMA/RoCE, optics, scale-out topology) and backend networks (storage fabric, management, out-of-band connectivity). Provide field feedback as GPU generations and interconnect standards evolve. Requirements: - 6+ years of experience in network engineering or data center deployment, including 2+ years managing engineers in a large-scale infrastructure environment. - Proven people leadership: Track record of building and retaining a high-performing technical team. - Program execution: Demonstrated ability to deliver on multiple concurrent, complex projects spanning sites, time zones, and vendors, with solid reporting and risk management. - Strong technical fluency: Solid grounding in physical layer standards (structured cabling, optics, 400G/800G and beyond), leaf-spine fabric (Arista/Juniper/NVIDIA-Mellanox), and protocols (BGP, EVPN-VXLAN, LLDP). Sufficient depth to guide your team and engage credibly with Architecture. - Automation fluency: Working understanding of Python/Ansible/ZTP-based deployment automation and CI/CD patterns, sufficient to drive adoption within your team. - Vendor and partner management: Experience holding external partners to SLAs, including escalation when performance slips. - Clear communication: Comfortable setting context and managing expectations with stakeholders across Ops, Supply Chain, and Architecture. - Operational judgment: Able to balance near-term delivery pressure against team health and sustainability. - Education: Bachelor's degree in a technical field or equivalent practical experience in hyperscale or ISP environments.

Similar roles