SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 300,000 / annual
Prime Intellect is building the open superintelligence stack—infrastructure that frontier AI labs develop internally, now available to every ambitious AI team. The company's platform, Lab, unifies compute, environments, evaluations, secure sandboxes, high-performance training, and deployment into one full-stack system for post-training at frontier scale. Prime Intellect has raised $150M from Founders Fund, Radical Ventures, NVIDIA, and exceptional operators including Andrej Karpathy and leaders from Ramp, Perplexity, Harvey, Datadog, Cognition, OpenAI, and others.
In this role, you will own the storage systems that feed frontier AI workloads. You'll design and operate storage architectures for training datasets, checkpointing, inference artifacts, and shared research workflows. Your responsibilities include deploying and tuning parallel filesystems, object storage, and local NVMe caching for demanding AI workloads; benchmarking throughput, latency, metadata performance, and concurrent access with representative training and checkpoint workloads; building provisioning, capacity planning, lifecycle management, and operational automation for storage services; designing and testing replication, recovery, backup, and failure-handling procedures with explicit durability and availability targets; diagnosing performance and reliability issues across applications, clients, networks, filesystems, and devices; and implementing access controls, tenant separation, quotas, monitoring, and runbooks while collaborating with compute and networking teams.
You'll work directly with customers pushing the boundaries of AI, from startups training foundation models to enterprises deploying massive inference infrastructure. You'll collaborate with a world-class engineering team while having direct impact on systems powering the next generation of AI breakthroughs.
REQUIREMENTS:
Required Experience:
- 3+ years building or operating production distributed storage systems
- Hands-on experience with at least one parallel or distributed filesystem or object storage platform (Lustre, BeeGFS, Ceph, or GPFS)
- Strong Linux administration and performance troubleshooting skills
- Experience automating infrastructure operations in Python, Go, Bash, or similar languages
- Understanding of storage failure modes, data integrity, consistency, replication, and recovery
Infrastructure Skills:
- Block, file, and object storage semantics and their performance tradeoffs
- NVMe/SSD performance, filesystem tuning, I/O profiling, and benchmarking
- High-throughput storage networking and distributed client behavior
- Capacity forecasting, observability, alerting, and safe maintenance procedures
- Authentication, authorization, encryption, and secure data lifecycle management
Nice to Have:
- Experience supporting large GPU training clusters and high-volume checkpoint workloads
- S3-compatible object storage, data tiering, or distributed caching
- RDMA-enabled storage or GPUDirect Storage experience
- Kubernetes storage integrations or SLURM environments
- Storage cost optimization and contributions to open-source storage systems