SlipstreamJobsFresh Startup & VC-Backed Jobs

Principal Systems Software Engineer

Crusoe - San Francisco, CA, United States - In-office - posted 2026-08-05

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Crusoe is building vertically integrated AI infrastructure, owning the full stack from energy generation to AI workloads. As Principal Systems Software Engineer, you will serve as the visionary technical lead for the next-generation AI infrastructure platform, bridging silicon and software design. You will architect and unify three core infrastructure pillars: Bare-Metal-as-a-Service (BMaaS) delivering raw GPU throughput via zero-latency InfiniBand/RDMA fabrics for massive-scale training; Intelligent IaaS with optimized thin virtualization layers (KVM or custom micro-VMs) providing enterprise isolation without virtualization overhead; and Elastic CaaS, a high-performance container substrate using Kubernetes or Slurm for AI workload bursting across heterogeneous GPU nodes. Key responsibilities include mastering the I/O path and leading architectural design of the internal cloud fabric, drawing on hyperscaler experience to drive the technical roadmap for SR-IOV, RDMA, and virtualized GPU scheduling. You will lead elite R&D workstreams to prototype and productionize novel methods for managing memory, networking, and compute not yet available in standard cloud distributions. You'll draft white papers and RFCs defining the next two years of Crusoe's compute and networking stack, work alongside Staff and Senior engineers resolving complex race conditions and optimizing kernel-level memory pinning for GPU clusters, and represent Crusoe in open-source communities influencing the global direction of cloud-native AI infrastructure. Required: 12+ years designing and shipping core infrastructure at major hyperscalers (OCI, AWS, Azure, GCP) or specialized HPC cloud; authoritative knowledge of Linux kernel, virtualization internals (KVM, QEMU, Firecracker), and high-performance networking (RoCE v2, InfiniBand); proven ability to design software maximizing NVIDIA/AMD GPU and high-speed NIC performance; experience leading cross-functional teams through high-ambiguity projects delivering production-ready systems; portfolio of significant contributions (patents, open-source, published research); rare ability to communicate technical nuances to both engineers and executives; Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent professional experience. Bonus: Patents in network virtualization, GPU scheduling, or distributed file systems; Linux Kernel, Kubernetes, or HPC project maintainer status; direct LLM training/inference infrastructure optimization experience; peer-reviewed publications at top systems venues (OSDI, SOSP, NSDI, SC); significant open-source contributions or industry white papers on distributed systems and high-performance networking.

Similar roles