SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Network Production Operations Engineer

Crusoe - San Francisco, CA, United States - In-office - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 165,000 - 200,000 / annual

Crusoe is a vertically integrated AI infrastructure company building the stack from electrons to tokens to power large-scale AI workloads. The company operates global network infrastructure including edge, backbone, data center fabric, and GPU cluster interconnects. You will serve as a Senior Network Production Operations Engineer supporting production reliability across Crusoe's hyperscale AI infrastructure. This is a hands-on role focused on incident response, root cause analysis, and automation-driven operational execution. Your work directly impacts the availability of AI workloads running across thousands of GPUs worldwide. Key responsibilities include: - Supporting uptime across Crusoe's global edge, backbone, data center, and GPU cluster network infrastructure - Writing and maintaining Python-based tooling to reduce operational toil and automate remediation workflows - Participating in high-severity network incident detection, triage, and mitigation - Performing root cause analysis using existing and newly built tooling - Building automation on top of monitoring stacks (streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, ThousandEyes) - Maintaining and improving runbooks, escalation playbooks, and standard operating procedures - Building and maintaining automated dashboards and alerting for reliability metrics in partnership with Architecture and SRE teams - Collaborating with Architecture and SRE teams on operational excellence initiatives You will work in a high-pressure environment where you take pride in keeping systems healthy and thrive on solving complex problems at scale. REQUIREMENTS: - 5+ years of production network engineering experience with focus on operations, incident response, and reliability in large-scale or internet-scale environments - Strong Python and scripting proficiency; comfortable writing diagnostic tooling and automation scripts - Hands-on experience with observability and monitoring tools: streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, ThousandEyes, and ability to script against their APIs - Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning - Solid hands-on knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments - Proficiency with Arista (EOS) and Juniper (Junos) platforms in leaf-spine CLOS architectures across multi-vendor environments - Comfort operating large device fleets across multi-region environments with on-call responsibility - Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience BONUS: - Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments - Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility - Exposure to contributing to SLIs and SLOs alongside SRE or product teams - Exposure to operating 10K+ device fleets across hyperscale or cloud environments

Similar roles