SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff Software Engineer, DC Infrastructure

Crusoe - San Francisco, CA, United States - In-office - posted 2026-08-12

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 250,000 - 300,000 / annual

Crusoe is a vertically integrated AI infrastructure company building the next generation of data center technology. The company owns and operates the full stack from power generation through GPU clusters, enabling efficient AI compute at scale. The Data Center Infrastructure Engineering (DCIE) team is responsible for the deployment, maintenance, observability, and automation of Crusoe's growing GPU fleet and data center operations. This Senior Staff Software Engineer role is a hands-on technical position focused on developing software for managing thousands of GPU servers and the facilities that house them. Key responsibilities include: - Developing deep-level diagnostics and troubleshooting tools for hardware faults in GPU racks and high-density compute systems - Building automation and AI agents for component-level diagnosis and remediation of failed or degraded hardware - Creating tooling for GPU platform management (NVIDIA A100, H200, GB200, B200, AMD 350X/355X) - Developing post-repair validation and testing tools using burn-in, PyTorch, and NVIDIA NCCL - Owning deployment, monitoring, and operational support of infrastructure tooling - Building automation for facilities management, power systems, and direct liquid cooling hardware - Collaborating with data center operations to develop innovative tooling and AI agents for critical environment management You will be a hands-on problem solver comfortable working independently while setting technical direction for specific projects. The role requires expertise in distributed systems, reliability, and cloud platforms (Kubernetes, Infrastructure as Code, GCP). Strong programming skills in Go, Python, Java, or Rust are essential. Experience with Temporal, hardware vendor relationships, or large-scale GPU fleet operations is valued. This is a high-impact position where you'll directly influence the scalability and performance of Crusoe's rapidly expanding GPU infrastructure, supporting the company's mission to solve the power bottleneck in AI compute.

Similar roles