SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Crusoe is building vertically integrated AI infrastructure, owning the full stack from energy generation to AI workloads. The company is solving the power bottleneck in AI compute through an energy-first approach.
As a Senior Hardware Systems Engineer, you will drive the end-to-end lifecycle of next-generation compute platforms, from prototype bring-up through large-scale production. You'll own performance characterization and validation strategies for CPU, GPU, and accelerated computing platforms, conducting in-depth workload characterization studies across training and inference workloads (dense, MoE, long-context, multimodal models) to understand compute, memory, communication, and I/O behavior.
Key responsibilities include translating workload and platform insights into cluster-level tuning and configuration recommendations (topology, parallelism strategy, scheduling, power, software stack settings) to maximize performance and efficiency. You'll build and maintain workload performance profiles and reference configurations that guide cluster deployment and scaling. You'll analyze system and workload performance to identify bottlenecks, lead complex system-level debugging across compute, memory, storage, networking, accelerators, and firmware, and partner with vendors and internal teams on prototyping and production readiness.
You'll collaborate across hardware, firmware, networking, software, infrastructure, reliability, and operations teams to resolve complex platform issues, using data and system-level insights to influence platform architecture, technology selection, and hardware roadmaps.
Required: 5-6+ years in hardware systems engineering, platform engineering, performance engineering, ML systems engineering, or infrastructure engineering. Hands-on experience with large-scale GPU/accelerated computing infrastructure for AI/ML or HPC workloads. Experience with distributed training/inference at scale, workload benchmarking, performance profiling, and system optimization. Strong understanding of modern server and accelerator architectures (CPU, GPU, memory, storage, networking, high-speed interconnects like PCIe, InfiniBand, NVLink). Experience with system bring-up, validation, performance characterization, and root-cause analysis. Ability to develop automation, testing, and diagnostics frameworks using Python or Shell. Strong analytical and problem-solving skills in ambiguous environments. Excellent technical communication and cross-functional collaboration. Bachelor's or Master's in Electrical Engineering, Computer Engineering, Computer Science, or equivalent.
Bonus: Experience influencing hardware/system configuration decisions based on workload performance data (HW/SW co-design). Deep experience with RDMA, RoCE, CXL, NVLink, or fabric-level performance analysis. Experience with inference serving frameworks, training frameworks, or ML compiler/runtime stacks. Familiarity with x86 and ARM-based server platforms. Experience building observability, diagnostics, or fleet-level performance and reliability systems.