SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Modal is building a new infrastructure layer for AI—a serverless cloud platform for AI, data, and compute-intensive applications. The company recently raised $355M in Series C funding at a $4.65B valuation and has crossed $300M+ ARR. Customers include Lovable, Ramp, Cognition, DoorDash, and Suno, who rely on Modal for instant GPU access, sub-second container starts, and native storage.
You will lead the team responsible for Modal's distributed object storage system—the backbone that underpins every container image, volume, and checkpoint on the platform. This system manages hundreds of petabytes of data replicated across multiple cloud object stores and a CDN, cached on local NVMe across a large fleet of workers in many datacenters, and shared peer-to-peer within each datacenter.
In this role, you will:
- Set technical direction for the storage primitives that other teams (filesystems, training, sandboxes) build on, balancing durability, latency, throughput, and cost
- Own the roadmap from today's hardest problems (garbage collection at petabyte scale, active-active replication, rate limiting) to architectural bets that shape the future (storage colocated with GPUs, tiered writes, capacity planning)
- Manage a team of 3–8 engineers while staying hands-on across the full stack: local disk and page cache, distributed blob storage, garbage collection, and observability
- Guide on-call practices and automation that keep the system healthy as it scales by orders of magnitude
- Participate in the on-call rotation and respond to production incidents
Key initiatives the team is working on include P2P data sharing across workers within a datacenter, replicating data across multiple blob storage providers, automating garbage collection at petabyte scale, and deploying colocated storage clusters to datacenters.
REQUIREMENTS:
- 7+ years of experience writing high-quality production code
- 3+ years of direct people management experience, ideally leading a team of engineers through project planning, growth, and performance conversations
- Experience building high-performance distributed storage or caching systems at large scale
- Strong cloud skills, including deep familiarity with object storage (S3 or similar), CDNs, and their consistency, throughput, and cost characteristics
- Strong knowledge of low-level operating system foundations (Linux kernel, file systems, page cache, containers)
- Experience with replication, content addressing, and consistency models in multi-region or multi-cloud systems
- Experience operating storage systems at scale (petabyte-scale datasets, high-throughput read/write paths, large-scale garbage collection or data migration), including owning cost and capacity planning
- Track record of setting technical direction and driving architectural decisions across a team, and of building primitives other teams depend on
- Willingness to step into the thick of it with on-call rotation and respond to production incidents
NICE-TO-HAVES:
- Experience with data engineering at petabyte scale
- Prior experience with Rust
About Modal Labs
AI / Data / Infrastructure — serverless cloud platform for AI, data, and compute-intensive applications.