SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 405,000 - 485,000 / annual
Anthropic is seeking a Senior Engineering Manager to lead the Capacity Engineering team, which manages one of the largest and fastest-growing infrastructure fleets in the industry. This team is responsible for ensuring all infrastructure resources are accounted for, well-utilized, and efficiently allocated across multiple accelerator families, CPU families, and cloud providers.
In this hands-on leadership role, you will lead a team of senior and staff-level engineers while staying close enough to the systems to review designs, make architectural decisions, and step into incidents when needed. You'll balance your time between people management, technical leadership, and cross-organizational alignment.
The team's work spans three overlapping areas:
**Data Platform**: Build and maintain pipelines that ingest occupancy and utilization telemetry from Kubernetes clusters, normalize billing and usage across cloud providers, and serve BigQuery tables queried by research engineers, finance, and leadership. This is product work as much as engineering.
**Planning and Assurance**: Make the state of the fleet legible and actionable in real time through cluster health tooling, capacity planning platforms, alerting on occupancy issues, and systemic fixes to scheduling and fragmentation.
**Efficiency**: Measure and improve how effectively every major workload uses its hardware across training, inference, and evals. Build benchmarking infrastructure and per-config baselines, then partner with system-owning teams to close performance gaps.
Key responsibilities include hiring and developing senior engineers, championing internal customers (research engineering, inference, infrastructure, finance teams), owning the engineering roadmap, setting technical standards, running the team as a product organization, partnering with cross-functional stakeholders, driving operational excellence, and scaling the function as the fleet diversifies.
**Requirements:**
- Experience managing software or infrastructure engineering teams, including hiring senior engineers, managing performance, and developing people into larger scope
- Strong technical background in production systems (data engineering, infrastructure, distributed systems, or observability) with hands-on experience you can draw on when reviewing designs or debugging
- Familiarity with at least one major cloud provider (AWS, GCP, or Azure), Kubernetes-based infrastructure, and modern observability stacks (e.g., Prometheus, Grafana)
- Track record of setting and executing an engineering roadmap in ambiguous, high-autonomy environments with many stakeholders and shifting priorities
- Excellent communication skills: ability to explain utilization metrics to research engineers and spend forecasts to CFOs, and advocate clearly for team priorities with senior leadership
- Comfort owning operational responsibility for systems the company depends on, including on-call and incident management
- Bachelor's degree or equivalent combination of education, training, and/or experience in a field relevant to the role
**Preferred Qualifications:**
- Experience leading teams working on capacity planning, resource management, product engineering, or FinOps at a hyperscaler or in a large-scale ML environment
- Familiarity with accelerator infrastructure (GPU metrics via DCGM, TPU utilization, or ML training/inference systems at the hardware level)
- Experience with multi-cloud billing and telemetry normalization (billing exports, reservation APIs, commitments, on-demand capacity reservations)
- Experience building or leading internal data products with self-service access, schema contracts, and documentation
- Background in scheduling, packing efficiency, or profiling-driven optimization of large distributed workloads