SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Network Production Engineer, Deployment

Crusoe - San Francisco, CA, United States - Hybrid - posted 2026-09-23

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 165,000 - 200,000 / annual

Crusoe is a vertically integrated AI infrastructure company building the future of energy-efficient AI compute. They own and operate every layer of the stack—from electrons to tokens—to power the world's most ambitious AI workloads. You'll join Crusoe Cloud as a Staff Network Production Engineer supporting the physical and logical implementation of their rapidly expanding global network of high-performance compute (HPC) and GPU-based AI infrastructure. This is a technical, hands-on role at the intersection of network engineering, data center operations, and project management. Key responsibilities: - Execute on-site and remote network deployments for new data center and edge site builds, from rack-and-stack through final cutover - Translate architecture designs into concrete deployment steps: cable maps, port assignments, device configs, and site-specific runbooks - Write and maintain Python/Ansible automation and ZTP workflows to stage, configure, and validate switches and routers at scale - Support burn-in and site acceptance testing (SAT) on new clusters, diagnose failures, and drive clean handoffs to Operations - Configure and troubleshoot Arista, Juniper, and NVIDIA/Mellanox gear in leaf-spine fabrics, including BGP, EVPN-VXLAN, and LLDP - Work directly with structured cabling vendors, remote hands, and data center providers to resolve physical layer issues - Diagnose physical and link-layer problems using OTDRs, light meters, and packet captures - Track hardware inventory and support turn-up of new backbone and edge interconnect capacity - Collaborate with other engineers on complex sites and flag recurring issues for process/automation improvements - Participate in on-call rotation for production network incidents Requirements: - 5+ years of network engineering experience with focus on large-scale data center deployments and infrastructure projects - Strong knowledge of physical layer standards: structured cabling (SMF/MMF, MPO/MTP), optical transceivers (400G/800G), data center power/cooling - Solid routing and switching knowledge: hands-on experience with Arista (EOS), Juniper (Junos), and NVIDIA/Mellanox in leaf-spine architecture - Working understanding of BGP, EVPN-VXLAN, and LLDP as they relate to large-scale fabric provisioning - Proficiency in Python and Ansible for automating deployment tasks and validating configuration state - Ability to manage multiple projects simultaneously across different time zones and physical locations - Strong troubleshooting skills for physical layer and link-layer issues - Bachelor's degree in a technical field or equivalent practical experience in hyperscale or ISP environments Bonus qualifications: - Experience in hyperscale, ISP, or large multi-tenant data center environments - Familiarity with GPU cluster networking (RDMA/RoCE, InfiniBand, NVIDIA NCCL-aware fabric design) - Exposure to network monitoring and observability tooling (Prometheus, Grafana, telemetry-based fabric health checks) - Vendor certifications (CCNP, JNCIP, Arista ACE) - Prior experience mentoring junior engineers or contributing to deployment standards/documentation

Similar roles