SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Baseten powers mission-critical inference for leading AI companies including Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma, and Writer. The company recently raised $1.5B in Series F funding and is building the platform engineers use to ship AI products at scale.
As Global Capacity Manager focused on TPUs, you will lead infrastructure strategy for Baseten's non-NVIDIA accelerator fleet, architecting and optimizing Google Cloud TPU capacity that powers customer AI workloads. You'll own the end-to-end capacity management journey—from securing large-scale TPU pod allocations to building automation ensuring reliable uptime across multi-cloud environments.
This is a hands-on engineering role bridging high-finance asset management and deep infrastructure engineering. You will act as fleet orchestrator for Google's TPU architecture, ensuring zero capacity outages while maintaining elite unit economics as Baseten diversifies beyond NVIDIA.
Key initiatives include: architecting TPU cluster infrastructure and deployment strategy with pod slicing and topology planning; building multi-cloud capacity management systems to move workloads seamlessly across TPU regions and pod configurations; developing automated operators to identify, cordon, and repair unhealthy TPU pods within an hour; and partnering with leadership and Google Cloud to secure dedicated TPU capacity for enterprise customers.
Responsibilities: lead TPU pod fleets managing full lifecycle of acquisition, allocation, and maintenance; execute complex workload migrations and deployment drains across TPU topologies with strict regional and compliance requirements; design and implement next-generation capacity management systems handling significant TPU volume growth alongside existing GPU fleet; build ROI models comparing TPU, GPU, and other accelerator options for profitable scaling; partner with Model Performance, SRE, Infra, and FDE teams ensuring workloads are tuned for TPU execution; lead capacity-crunch incident response by rapidly reallocating and re-coordinating TPU workloads.
Requirements: Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Mathematics, or related field; 5+ years professional experience in high-growth environment, preferably at hyperscaler (GCP, AWS, Azure) or specialized accelerator provider; hands-on experience with Google Cloud TPUs including pod slicing, ICI topology, JAX/XLA, and TPU-specific scheduling and fault handling; deep Kubernetes expertise including taints, cordons, node draining, and custom operators; production-level experience with Go or Python; strong financial literacy and ability to model complex trade-offs between capacity reliability and cost; high tenacity and collaborative mindset.
Nice to have: experience with additional non-NVIDIA accelerators (AWS Trainium/Inferentia, AMD Instinct); familiarity with multi-accelerator scheduling and cost/performance tradeoff modeling; prior experience optimizing workloads with model performance or ML systems teams.