SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: GBP 325,000 - 485,000 / annual
Anthropic is building reliable, interpretable, and steerable AI systems. The Kubernetes Platform team owns and operates the control plane for one of the industry's largest AI compute fleets, spanning multiple cloud providers and datacenters used to train, research, and serve frontier AI models.
This is a Staff+ individual contributor role on a team operating at unprecedented scale. The platform runs thousands of accelerators across Kubernetes clusters where traditional defaults no longer work. You will own critical infrastructure components that directly enable Anthropic's ability to reliably and safely train frontier models as compute footprint grows exponentially.
Key responsibilities include:
- Own, operate, and extend the Kubernetes scheduler for accelerator fleets, including custom scheduling plugins for gang scheduling, topology awareness, and preemption
- Scale the Kubernetes control plane (apiserver, etcd, controller-manager) far beyond typical limits and identify bottlenecks proactively
- Design and operate core cluster services like service discovery that every workload depends on
- Build and maintain custom controllers, operators, and CRDs
- Partner with research, training, and inference teams to translate workload requirements into platform capabilities
- Collaborate with cloud providers on required features and escalations
- Participate in on-call, lead incident response, and design reliability processes (postmortems, runbooks, SLOs)
Required qualifications:
- Significant production experience building and operating distributed systems
- Proficiency in systems languages (Go, Python, Rust, C++)
- Deep, hands-on Kubernetes expertise beyond user-level, including scheduler, controllers, apiserver, or large multi-tenant cluster operations
- Ability to debug complex issues across the full stack from API behavior to node and network root causes
- Track record designing for reliability, correctness, and clear failure semantics in systems others depend on
- Strong written and verbal communication with ability to build consensus
Preferred: Kubernetes internals contributions, cluster scheduler or batch system experience (Kueue, Volcano, Slurm), control plane scaling background, ML infrastructure familiarity (GPUs, TPUs, gang scheduling, NCCL), GCP/AWS experience, low-level systems knowledge (Linux kernel, cgroups, eBPF), 12+ years relevant experience including leading large infrastructure projects.