SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 170,000 - 205,000 / annual
Crusoe is building vertically integrated AI infrastructure, owning the full stack from energy generation to AI workloads. The company is solving the compute-power bottleneck with an energy-first approach to sustainable AI infrastructure.
As a Senior Hardware Systems Engineer, you will own the complete hardware lifecycle for Crusoe's GPU and CPU-based compute systems—from prototype bring-up through large-scale production. You'll drive automation, deep issue resolution, and reliability improvements across Crusoe Cloud's infrastructure, with particular focus on PCIe, InfiniBand, and NVMe/storage subsystems.
Key responsibilities include:
- Leading the full hardware development and sustaining lifecycle, including feasibility studies, bring-up, validation, deployment, and production support
- Developing and maintaining scripting and automation frameworks for hardware testing, diagnostics, and continuous reliability improvements
- Conducting deep troubleshooting and debugging across PCIe (link training, topology, performance), InfiniBand (fabric debugging, throughput, connectivity), and NVMe/storage (performance bottlenecks, firmware interactions, failure analysis)
- Performing rigorous system validation and characterization for GPU, CPU, and high-performance compute platforms
- Supporting end-to-end integration and solution testing to ensure products meet performance, reliability, and scalability expectations
- Collaborating with mechanical, thermal, firmware, software, and manufacturing teams to resolve system-level issues
- Driving prototyping, qualification, and readiness for high-volume manufacturing with internal teams and external vendors
- Identifying opportunities for new hardware technologies, testing methods, and sustainability improvements
- Providing data-driven insights to influence hardware roadmap and reliability strategy
You bring 8–10+ years of hardware development, validation, sustaining engineering, or production engineering experience. You have strong hands-on expertise in PCIe, InfiniBand, and NVMe/storage debugging; deep proficiency in hardware bring-up, board-level debugging, and system-level validation; and the ability to design and implement automation frameworks using Python, Shell, or similar languages. You have a technical background in digital and analog design, server architecture, and high-performance compute hardware, with experience working across thermal, mechanical, firmware, and software functions. A Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, or equivalent experience is required.
Bonus qualifications include experience designing or optimizing GPU-to-GPU communication architectures for AI/ML workloads, direct experience integrating NVLink or next-generation GPU interconnect technologies, familiarity with cutting-edge GPU architectures, expertise supporting systems across ARM and x86 server architectures, background in sustainable or energy-efficient hardware design, and advanced certifications in AI/HPC hardware systems.