SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 320,000 - 485,000 / annual
Anthropic's Inference team is seeking a Staff or Senior Software Engineer to design, build, and maintain the distributed systems that serve Claude to millions of users worldwide. This role sits at the intersection of infrastructure excellence and AI research enablement.
You will own critical components of Anthropic's inference stack, including intelligent request routing, load balancing, and fleet orchestration across thousands of AI accelerators spanning multiple cloud providers (AWS, GCP, Azure). The team balances two competing mandates: maximizing compute efficiency to support explosive customer growth while providing high-performance infrastructure that enables researchers to develop next-generation models.
Key responsibilities include:
- Designing and building resilient distributed systems that serve hundreds of thousands of customers daily
- Developing intelligent request routing, load balancing, and traffic management across diverse hardware and cloud platforms
- Optimizing compute efficiency and cost through autoscaling and workload orchestration across production, research, and experimental environments
- Building production-grade deployment pipelines for reliable model releases
- Integrating new AI accelerator platforms and supporting inference for emerging model architectures
- Analyzing observability data to tune performance based on real-world production workloads
- Managing multi-region deployments and geographic routing for global customers
You should have significant experience with distributed systems, ideally including high-performance, large-scale deployments. Preferred experience includes machine learning systems at scale, load balancing/routing systems, LLM inference optimization, Kubernetes, and proficiency in Python or Rust. This is a hands-on technical role where you'll directly impact both business results and research breakthroughs.