SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Together AI is seeking a Staff Engineer to architect, build, and operate multi-petabyte distributed storage systems purpose-built for the world's largest AI training and inference workloads. You will take ownership of the technical strategy and storage roadmap, driving high-performance architectural decisions as the company scales its GPU fleet.
Key responsibilities include:
- Architecting and implementing storage strategy and roadmap for AI/ML workloads at massive scale
- Engineering and scaling multi-petabyte storage systems by integrating technologies like Vast, Weka, and Ceph, with deep cost optimization through automated tiering and lifecycle policies
- Developing intelligent caching and tiered storage architectures to achieve extreme IOPS and cluster-wide throughput for GPU-scale training and inference
- Tuning storage isolation at network layers to ensure secure, production-grade multi-tenancy
- Building Kubernetes-native storage operators and controllers for automated provisioning, self-service abstractions, and quota enforcement
- Engineering end-to-end data paths to achieve 10+ GB/s per GPU node, architecting multi-tier caching for model weights and datasets, and tuning parallel filesystems using advanced profiling
- Optimizing data paths through benchmarking and profiling, contributing to open-source storage projects and internal tooling
You will operate at extreme scale: managing multi-petabyte systems, optimizing for 10-50 GB/s per node throughput, designing for 99.999%+ uptime, and scaling infrastructure across thousands of nodes. This is a hands-on technical leadership role where you drive system design decisions that significantly improve performance, reliability, and cost efficiency.
Together AI is an AI-native cloud platform trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, serving 400+ trillion tokens monthly. The platform enables AI engineers to run, fine-tune, and pre-train open-source and custom models at scale.
REQUIREMENTS:
- 8+ years of storage engineering experience, managing distributed storage at multi-petabyte scale
- Proven track record deploying and operating high-performance storage for GPU/HPC clusters
- Deep Kubernetes and cloud-native storage experience in production environments
- Strong coding skills in Go and Python with demonstrated ability to build production-grade systems and tooling
- BS/MS in Computer Science, Engineering, or equivalent practical experience
- History of technical leadership: designing systems that significantly improved performance, reliability (99.999%+ uptime), or cost efficiency
- Deep expertise in at least one of: Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale
- Production experience with S3, MinIO, Ceph, or R2 including performance optimization and cost management
- Kubernetes storage expertise: CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers
- Storage optimization for GPU workloads, RDMA/InfiniBand networking, parallel filesystem optimization (TB/s aggregate cluster throughput)
- Infrastructure as Code: Terraform, Ansible, Helm, GitOps (ArgoCD)
- Advanced Linux storage stack knowledge: filesystems (ext4, xfs), LVM, NVMe optimization, RAID configurations
- Observability tools: Prometheus, Grafana, Thanos architecture and operations
NICE TO HAVE:
- GPU Direct Storage (GDS), NVMe-oF, storage networking, RDMA implementations
- ML/AI storage patterns (model weights, checkpointing, dataset caching).
- Storage benchmarking and profiling tools (fio, iperf3, iostat, blktrace)
About Together AI
AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.