SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Engineer, Distributed Storage and HPC & AI Infrastructure

Together AI - Bangalore, KA, India - In-office - posted 2026-09-11

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Together AI is seeking a Staff Engineer to architect, build, and operate multi-petabyte distributed storage systems purpose-built for the world's largest AI training and inference workloads. You will take ownership of the technical strategy and storage roadmap, driving high-performance architectural decisions as the company scales its GPU fleet. Key responsibilities include: - Architecting and implementing storage strategy and roadmap for AI/ML workloads at massive scale - Engineering and scaling multi-petabyte storage systems by integrating technologies like Vast, Weka, and Ceph, with deep cost optimization through automated tiering and lifecycle policies - Developing intelligent caching and tiered storage architectures to achieve extreme IOPS and cluster-wide throughput for GPU-scale training and inference - Tuning storage isolation at network layers to ensure secure, production-grade multi-tenancy - Building Kubernetes-native storage operators and controllers for automated provisioning, self-service abstractions, and quota enforcement - Engineering end-to-end data paths to achieve 10+ GB/s per GPU node, architecting multi-tier caching for model weights and datasets, and tuning parallel filesystems using advanced profiling - Optimizing data paths through benchmarking and profiling, contributing to open-source storage projects and internal tooling You will operate at extreme scale: managing multi-petabyte systems, optimizing for 10-50 GB/s per node throughput, designing for 99.999%+ uptime, and scaling infrastructure across thousands of nodes. This is a hands-on technical leadership role where you drive system design decisions that significantly improve performance, reliability, and cost efficiency. Together AI is an AI-native cloud platform trusted by Cursor, Decagon, ElevenLabs, Salesforce, and Zoom, serving 400+ trillion tokens monthly. The platform enables AI engineers to run, fine-tune, and pre-train open-source and custom models at scale. REQUIREMENTS: - 8+ years of storage engineering experience, managing distributed storage at multi-petabyte scale - Proven track record deploying and operating high-performance storage for GPU/HPC clusters - Deep Kubernetes and cloud-native storage experience in production environments - Strong coding skills in Go and Python with demonstrated ability to build production-grade systems and tooling - BS/MS in Computer Science, Engineering, or equivalent practical experience - History of technical leadership: designing systems that significantly improved performance, reliability (99.999%+ uptime), or cost efficiency - Deep expertise in at least one of: Ceph, WekaFS, Lustre, Vast, GPFS, or similar parallel filesystems at multi-petabyte scale - Production experience with S3, MinIO, Ceph, or R2 including performance optimization and cost management - Kubernetes storage expertise: CSI drivers, StatefulSets, PersistentVolumes, storage operators, and custom controllers - Storage optimization for GPU workloads, RDMA/InfiniBand networking, parallel filesystem optimization (TB/s aggregate cluster throughput) - Infrastructure as Code: Terraform, Ansible, Helm, GitOps (ArgoCD) - Advanced Linux storage stack knowledge: filesystems (ext4, xfs), LVM, NVMe optimization, RAID configurations - Observability tools: Prometheus, Grafana, Thanos architecture and operations NICE TO HAVE: - GPU Direct Storage (GDS), NVMe-oF, storage networking, RDMA implementations - ML/AI storage patterns (model weights, checkpointing, dataset caching). - Storage benchmarking and profiling tools (fio, iperf3, iostat, blktrace)

About Together AI

AI / Data / Infrastructure — cloud platform for open-source and generative AI model training and inference.

Similar roles