SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, DC Infrastructure

Crusoe - San Francisco, CA, United States - In-office - posted 2026-08-12

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 215,000 - 260,000 / annual

Crusoe is a vertically integrated AI infrastructure company building the next generation of data center systems to power AI workloads. The company owns and operates the full stack from energy generation through GPU deployment, with a mission to solve the power bottleneck constraining AI compute availability. The Data Center Infrastructure Engineering (DCIE) team is responsible for the design, deployment, maintenance, observability, and automation of Crusoe's GPU fleet and data center operations. This Staff Software Engineer role is a hands-on, high-impact position focused on developing software systems that manage and maintain thousands of GPU servers across multiple data centers. Key responsibilities include: - Designing and implementing deep-level diagnostics and troubleshooting tools for hardware faults in GPU racks and high-density compute systems - Building automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware - Developing observability and monitoring tooling for GPU platforms (NVIDIA A100, H200, GB200, B200, AMD 350X/355X) - Creating post-repair validation and testing tools using burn-in, PyTorch, and NVIDIA NCCL - Owning deployment, monitoring, and operational support of infrastructure tooling to maximize GPU fleet availability - Collaborating with data center operations on critical environment management and facilities automation - Setting technical direction for specific projects and executing independently Required qualifications include strong software engineering fundamentals, expertise in distributed systems and cloud platforms (Kubernetes, IaC, GCP), proficiency in at least one of Go, Python, Java, or Rust, and demonstrated ability to identify problems, develop scalable solutions, and ship rapidly. Nice-to-have skills include Temporal experience, direct hardware vendor relationships, and background in hyperscale GPU fleet or data center operations. This is a unique opportunity to work on infrastructure that directly enables the AI revolution, with exposure to cutting-edge hardware, distributed systems at scale, and the operational challenges of managing some of the world's most demanding compute environments.

Similar roles