SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anthropic is seeking an Engineering Manager to lead the Inference Infrastructure team, responsible for the control plane that coordinates Claude's inference fleet across all serving surfaces (claude.ai, API, cloud partners, internal research). This deeply technical group designs placement and load-balancing algorithms, builds quantitative models of demand and system performance, optimizes latency across kernel and network boundaries, and manages the inference request path end-to-end.
You will lead a strong team of ML platform, infrastructure, and distributed-systems engineers, working alongside teams building ML internals and cloud infrastructure. The role requires deep systems expertise to make architectural decisions about fleet coordination, evaluate candidates with kernel-level knowledge, and understand how changes ripple across the entire system.
Key responsibilities include: owning the technical roadmap for inference fleet coordination (traffic routing, capacity placement, caching, demand response, control-plane protocols); partnering with product, inference engine, performance, and capacity teams to identify and ship throughput, latency, utilization, and cost improvements with measurable results; building a culture of quantitative modeling and evidence-based decision-making; setting technical strategy for heterogeneous hardware and multi-cloud evolution; running operational backbone (on-call, incident response, postmortems, deploy safety); creating clarity at the seam between API surface, inference engines, capacity planning, and cloud deployment; developing and retaining strong teams while hiring to a high technical bar; coaching engineers through shifting priorities; shaping team structure as scope grows; and unblocking critical initiatives when needed.
Minimum qualifications: engineering management experience leading critical-path production infrastructure at scale; deep systems background (load balancing, scheduling, cluster orchestration, autoscaling, distributed state, high-performance networking); proven track record shipping performance/efficiency improvements in large-scale systems with quantified impact; production infrastructure operations experience (on-call, incident response, capacity events, deploy discipline); results-oriented, impact-driven approach; ability to build cross-team relationships; curiosity about ML systems and transformer inference.
Preferred: 5+ years engineering management; LLM inference serving experience (KV caching, continuous batching, request scheduling, prefill/decode disaggregation); background in cluster schedulers, autoscalers, load balancers, service meshes, or fleet control planes (Kubernetes, Borg-style systems); multi-cloud or partner platform experience.