SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Production Engineer, SDN

Crusoe - Sunnyvale, CA, United States - In-office - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 170,000 - 205,000 / annual

Crusoe is a vertically integrated AI infrastructure company building sustainable cloud platforms powered by an energy-first approach. The Production Engineering team maintains performance and reliability of AI-optimized cloud infrastructure at scale. In this Senior PE role focused on software-defined networking (SDN), you will own the availability, performance, and scalability of Crusoe's cloud networking services that power compute-intensive, latency-sensitive AI and HPC workloads. You'll build automation and self-healing tools to monitor and maintain SDN infrastructure spanning edge, backbone, and data center connectivity. Key responsibilities include: - Designing and implementing automation for network provisioning, configuration, and remediation across distributed systems - Driving reliability initiatives around routing convergence, traffic engineering, policy enforcement, and fault isolation - Supporting high-performance virtualized networking services (overlays, service meshes, programmable data planes) for large-scale AI compute clusters - Investigating and resolving networking incidents using deep telemetry, flow logs, and packet-level analysis - Optimizing network stacks, drivers, and control planes in partnership with hardware and kernel teams - Contributing to fault-tolerant, scalable SDN backend architecture for AI-first environments - Maintaining user-facing networking services with focus on availability, latency reduction, and error budget adherence You bring 5+ years of professional PE, systems, or networking engineering experience with demonstrated expertise in SDN platforms (Open vSwitch, ONOS, Tungsten Fabric, Calico, Cilium, EVPN/VXLAN). You have hands-on proficiency in Python, Go, or C; deep Linux networking knowledge; strong containerization and Kubernetes CNI experience; and familiarity with managed networking services at scale (AWS VPC, GCP VPC, Azure Virtual Network). You excel at incident response, troubleshooting, and cross-functional collaboration.

Similar roles