SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 285,000 - 335,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the energy and software stack to power the world's most ambitious AI workloads. As Director of Engineering for Flex Compute, you'll lead the design and delivery of a utility-facing, safety-critical control system that manages dynamic power allocation across Crusoe's data center fleet.
You'll own the curtailment orchestration layer—the system that reduces data-center power draw on demand while protecting critical workloads. This includes grid-signal ingestion, staged load shedding, per-SKU power capping, and dynamic power management for GPU oversubscription. You'll land utility-validated pilots with tight accuracy and high-fidelity telemetry, then scale to production.
Key responsibilities include designing safety-critical systems with authenticated signal ingress and fail-safe defaults; building workload-aware GPU power estimation validated against fleet telemetry; integrating with on-site battery and generation systems; and partnering deeply with Data Center Engineering and Energy teams on interconnection commitments and curtailment program design. You'll also integrate with Crusoe's cloud control plane across Kubernetes and Slurm fleets, make build-vs-leverage decisions on vendor power-management stacks, and hire and lead the engineering team from the ground up.
You bring 12+ years of software engineering experience with 5+ years leading engineering teams, ideally taking systems from 0→1 to production scale. You have deep experience with distributed control planes, orchestration, or fleet automation (Temporal, Kubernetes, Slurm); a track record with safety-critical or physically-actuating systems; and working knowledge of data-center power systems, utility interconnection, switchgear, UPS, BESS, rack/PDU distribution, and GPU power management. Experience with energy markets, grid programs, demand response, or ISO/RTO market signals is valuable. Bonus experience includes GPU cluster operations, AI cloud infrastructure, energy sector background, GPU power/performance modeling, or checkpoint/preemption for large training jobs.