SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer II (DCIE)

Crusoe - San Francisco, CA, United States - In-office - posted 2026-08-04

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 140,000 - 165,000 / annual

Crusoe is a vertically integrated AI infrastructure company building the energy and compute backbone for the world's most ambitious AI workloads. The company owns and operates the full stack—from power generation through GPU fleet management—to solve the critical bottleneck of AI compute availability. You'll join the Data Center Infrastructure Engineering team as a Software Engineer II, responsible for developing and maintaining software systems that manage Crusoe's rapidly expanding fleet of GPU servers and the data centers housing them. This is a hands-on role focused on building diagnostic, observability, automation, and repair tooling for high-performance GPU compute clusters. Key responsibilities include: - Developing deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems - Building automation and troubleshooting tools for NVIDIA (A100, H200, GB200, B200) and AMD (350X/355X) GPU platforms - Creating AI agents for component-level diagnosis and hardware remediation - Collaborating with data center operations to develop innovative tooling for critical environment management - Developing post-repair validation and testing tools (burn-in, PyTorch, NVIDIA NCCL) to ensure system stability - Owning deployment, monitoring, and operational support of developed solutions to maximize GPU fleet availability - Building automation for facilities management, power systems, and direct liquid cooling hardware You bring 2-3 years of software engineering experience with demonstrated ability to identify problems, rapidly develop scalable solutions, and ship them. You're proficient in at least one of Go, Python, Java, or Rust, with expertise in distributed systems, reliability, and cloud platforms (Kubernetes, IaC, GCP). You're a strong analytical problem-solver with excellent communication skills, comfortable working independently and collaborating with senior engineers on critical initiatives. Nice-to-have qualifications include experience with Temporal and Kubernetes, direct work with hardware vendors, or background in large-scale GPU fleet or hyperscale data center operations.

Similar roles