SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 250,000 - 300,000 / annual
Crusoe is a vertically integrated AI infrastructure company building the complete stack from energy to tokens to power large-scale AI workloads. As a Senior Staff Deployment Automation Engineer on the Compute Team, you will own deployment and testing automation for large-scale, multi-node GPU clusters across Crusoe's AI Cloud.
You will be responsible for the CI/CD infrastructure that enables rapid, reliable releases and deployments across datacenters. Key responsibilities include: designing and executing deployment automation for bare-metal, on-premise systems; building CI/CD platforms using GitLab and modern tooling; creating large-scale validation tests for multi-node virtualized GPU clusters; maintaining Linux configurations with Ansible, AWX, and observability tools; orchestrating canary deployments and blue-green testing with automatic rollback; developing automation frameworks in Python or Go to provision and stress-test environments; and creating automated test suites using fio, stress-ng, and iperf.
You will need 12+ years of experience building and deploying automated integration testing for AI cloud environments, from low-level Linux systems to distributed control planes. Required skills include deep knowledge of Kubernetes, Docker, Terraform, and Postgres; intimate familiarity with CI/CD pipelines and GitLab; experience with configuration management systems (Ansible, Puppet, Chef, SaltStack); advanced Python and Bash proficiency; Linux kernel internals (PCIe, VFIO, memory management); and familiarity with NVIDIA CUDA/NCCL or AMD ROCm/RCCL in multi-node contexts. Strong understanding of RDMA, RoCE, and InfiniBand protocols is essential.
Bonus experience includes MNNVL (Multi-Node NVLink), hardware debugging tools (NVIDIA Nsight, AMD Omniperf), and containerized GPU orchestration with Kubernetes. This is a hands-on technical leadership role where you'll drive infrastructure stability and enable teams across the organization to deploy reliably at scale.