SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 165,000 - 200,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the stack from electrons to tokens to power large-scale AI workloads. The company operates global network infrastructure including edge, backbone, data center fabric, and GPU cluster interconnects.
You will serve as a Senior Network Production Operations Engineer supporting production reliability across Crusoe's hyperscale AI infrastructure. This is a hands-on role focused on incident response, root cause analysis, and automation-driven operational execution. Your work directly impacts the availability of AI workloads running across thousands of GPUs worldwide.
Key responsibilities include:
- Supporting uptime across Crusoe's global edge, backbone, data center, and GPU cluster network infrastructure
- Writing and maintaining Python-based tooling to reduce operational toil and automate remediation workflows
- Participating in high-severity network incident detection, triage, and mitigation
- Performing root cause analysis using existing and newly built tooling
- Building automation on top of monitoring stacks (streaming telemetry, SNMP, NetFlow, Kentik, Grafana, Prometheus, ThousandEyes)
- Maintaining and improving runbooks, escalation playbooks, and standard operating procedures
- Building and maintaining automated dashboards and alerting for reliability metrics in partnership with Architecture and SRE teams
- Collaborating with Architecture and SRE teams on operational excellence initiatives
You will work in a high-pressure environment where you take pride in keeping systems healthy and thrive on solving complex problems at scale.
REQUIREMENTS:
- 5+ years of production network engineering experience with focus on operations, incident response, and reliability in large-scale or internet-scale environments
- Strong Python and scripting proficiency; comfortable writing diagnostic tooling and automation scripts
- Hands-on experience with observability and monitoring tools: streaming telemetry, SNMP, NetFlow/sFlow, Grafana, Prometheus, ThousandEyes, and ability to script against their APIs
- Experience operating RDMA/RoCE lossless fabrics for GPU or HPC workloads, including PFC, ECN, and DCQCN tuning
- Solid hands-on knowledge of BGP, EVPN-VXLAN, IS-IS, OSPF, MPLS, QoS, and TCP/IP in production data center environments
- Proficiency with Arista (EOS) and Juniper (Junos) platforms in leaf-spine CLOS architectures across multi-vendor environments
- Comfort operating large device fleets across multi-region environments with on-call responsibility
- Bachelor's degree in Computer Science, Electrical Engineering, or related field, or equivalent practical experience
BONUS:
- Experience with NVIDIA/Mellanox networking platforms in GPU cluster environments
- Familiarity with Kentik or Arbor for traffic analysis and DDoS visibility
- Exposure to contributing to SLIs and SLOs alongside SRE or product teams
- Exposure to operating 10K+ device fleets across hyperscale or cloud environments