SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 170,000 - 205,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the energy and compute layer for the world's most ambitious AI workloads. The company owns and operates the full stack from power generation through GPU fleet management, addressing the critical bottleneck of energy availability for AI compute.
You will join the Data Center Infrastructure Engineering team as a Senior Software Engineer focused on GPU fleet management and data center operations. This is a hands-on role developing diagnostic, observability, automation, and repair tooling for high-performance GPU compute clusters at scale.
Key responsibilities include:
- Developing deep-level diagnostics and troubleshooting for hardware faults in GPU racks and high-density compute systems
- Building troubleshooting and automation tooling for NVIDIA (A100, H200, GB200, B200) and AMD (350X, 355X) GPU platforms
- Creating AI agents for component-level diagnosis and hardware remediation
- Collaborating with data center operations to develop innovative tooling for critical environment management
- Developing post-repair validation and testing tools (burn-in, PyTorch, NVIDIA NCCL) to ensure system stability
- Owning deployment, monitoring, and operational support of developed solutions to maximize GPU fleet availability
- Building automation for facilities management, power systems, and direct liquid cooling hardware
You bring 4-6 years of software engineering experience with demonstrated ability to identify problems, rapidly develop scalable solutions, and ship them. You have expertise in distributed systems, reliability, and cloud platforms (Kubernetes, Infrastructure-as-Code, GCP). You're proficient in at least one of Go, Python, Java, or Rust, with strong analytical and problem-solving skills. You work effectively both independently and collaboratively.
Nice-to-have qualifications include experience with Temporal and Kubernetes, direct work with hardware vendors, and background in large-scale GPU fleet operations or hyperscale data center environments.