SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 170,000 - 205,000 / annual
Crusoe is a vertically integrated AI infrastructure company building sustainable cloud platforms powered by an energy-first approach. The company owns and operates the full stack from power generation through AI services to support compute-intensive workloads.
As a Senior Production Engineer focused on Storage, you will be a core member of the Site Reliability Engineering team responsible for maintaining Crusoe's AI-optimized cloud storage infrastructure. This role is critical to ensuring availability, performance, and scalability of storage systems powering large-scale AI and HPC compute clusters.
Key responsibilities include:
- Building automation and self-healing tools to monitor and maintain distributed cloud storage infrastructure (block, file, and object storage)
- Driving reliability initiatives around data replication, encryption, backup/restore, and failover mechanisms
- Implementing and maintaining high-performance NVMe and SSD-backed volumes for AI compute clusters
- Supporting user-facing storage services with focus on availability, performance tuning, and error budgets
- Investigating and resolving storage incidents using deep telemetry, logs, and performance profiling
- Partnering with hardware and kernel teams to diagnose low-level I/O issues and optimize I/O paths, cache policies, and file systems
- Contributing to architecture of fault-tolerant, scalable storage backends for AI-first cloud environments
Required qualifications:
- Bachelor's degree in Computer Science, Electrical Engineering, or related field (or equivalent experience)
- 5+ years professional experience in Storage SRE, systems, or storage engineering
- Deep hands-on experience with enterprise storage platforms (Pure Storage, EMC)
- Deep understanding of object, block, and file storage paradigms
- Proficiency in Go, Python, Java, or C
- Experience with Infrastructure as Code (Terraform, Ansible, Puppet)
- Deep Linux internals knowledge (I/O subsystems, memory management, storage scheduling)
- Familiarity with storage protocols (NFS, SMB, iSCSI, NVMe-oF)
- Strong experience with containerized workloads and Kubernetes/Docker
- Excellent incident response and troubleshooting practices
- Experience operating managed storage services at scale (AWS, GCP, Azure)
Nice-to-have skills include distributed storage systems (Ceph, GlusterFS, OpenEBS), open-source storage contributions, and hybrid on-prem/cloud storage experience.