SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Production Engineer, Storage

Crusoe - Sunnyvale, CA, United States - In-office - posted 2026-07-31

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 170,000 - 205,000 / annual

Crusoe is a vertically integrated AI infrastructure company building sustainable cloud platforms powered by an energy-first approach. The company owns and operates the full stack from power generation through AI services to support compute-intensive workloads. As a Senior Production Engineer focused on Storage, you will be a core member of the Site Reliability Engineering team responsible for maintaining Crusoe's AI-optimized cloud storage infrastructure. This role is critical to ensuring availability, performance, and scalability of storage systems powering large-scale AI and HPC compute clusters. Key responsibilities include: - Building automation and self-healing tools to monitor and maintain distributed cloud storage infrastructure (block, file, and object storage) - Driving reliability initiatives around data replication, encryption, backup/restore, and failover mechanisms - Implementing and maintaining high-performance NVMe and SSD-backed volumes for AI compute clusters - Supporting user-facing storage services with focus on availability, performance tuning, and error budgets - Investigating and resolving storage incidents using deep telemetry, logs, and performance profiling - Partnering with hardware and kernel teams to diagnose low-level I/O issues and optimize I/O paths, cache policies, and file systems - Contributing to architecture of fault-tolerant, scalable storage backends for AI-first cloud environments Required qualifications: - Bachelor's degree in Computer Science, Electrical Engineering, or related field (or equivalent experience) - 5+ years professional experience in Storage SRE, systems, or storage engineering - Deep hands-on experience with enterprise storage platforms (Pure Storage, EMC) - Deep understanding of object, block, and file storage paradigms - Proficiency in Go, Python, Java, or C - Experience with Infrastructure as Code (Terraform, Ansible, Puppet) - Deep Linux internals knowledge (I/O subsystems, memory management, storage scheduling) - Familiarity with storage protocols (NFS, SMB, iSCSI, NVMe-oF) - Strong experience with containerized workloads and Kubernetes/Docker - Excellent incident response and troubleshooting practices - Experience operating managed storage services at scale (AWS, GCP, Azure) Nice-to-have skills include distributed storage systems (Ceph, GlusterFS, OpenEBS), open-source storage contributions, and hybrid on-prem/cloud storage experience.

Similar roles