SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 500,000 - 850,000 / annual
Anthropic is seeking a Research Engineer to work on distributed systems that power reinforcement learning at frontier scale. In this role, you'll design, build, and operate the infrastructure that trains Claude to reason, write code, and act autonomously over long horizons.
The RL distributed system is an unusually demanding platform where training, sampling, and environment execution run concurrently across large fleets of accelerators and hosts, exchanging data continuously while hardware fails and research priorities shift. Your work will directly determine how much compute translates into learning and how quickly the research team can iterate.
Key responsibilities include:
- Design and build distributed systems for training, sampling, and environment execution at scale
- Identify and remove bottlenecks in scheduling, data movement, storage, networking, and coordination
- Implement fault tolerance across all layers: failure detection, isolation, and recovery for long-running jobs
- Design resource management and autoscaling systems that adapt as workload demands shift
- Build observability and diagnostics tools to understand run behavior, performance degradation, and unexpected results
- Create automation for detecting and remediating common problems, with safe operational interfaces
- Collaborate with researchers and performance engineers to preserve training correctness and prevent subtle nondeterminism
- Conduct incident reviews and redesign systems to eliminate failure classes at their source
You'll work on representative projects such as designing heterogeneous cluster schedulers, building failure detection and recovery systems, scaling environment execution without increasing latency, designing dynamic autoscaling policies, creating diagnostics systems, debugging rare data corruption across services, and designing operational interfaces for automated tools.
Anthropics's mission is to create reliable, interpretable, and steerable AI systems. The team is collaborative and values communication skills, viewing AI research as an empirical science akin to physics and biology. The company operates as a single cohesive team focused on large-scale research efforts that advance steerable, trustworthy AI.
QUALIFICATIONS:
Minimum:
- Strong software engineering skills in Python and at least one systems language (Rust, C++, or Go)
- Experience designing, building, and operating large-scale distributed systems in production
- Deep understanding of distributed systems fundamentals: consistency, coordination, consensus, failure modes, and recovery
- Ability to reason quantitatively about throughput, latency, and resource costs across compute, memory, storage, and network
- Experience debugging complex failures across many hosts and services, including non-reproducible failures
- Strong written communication skills, including design documents and incident writeups
- Bachelor's degree or equivalent combination of education, training, and professional experience in a relevant field
Preferred:
- Experience running ML training or inference infrastructure at scale
- Experience across multiple stack layers (scheduling, storage, networking, orchestration)
- Experience building schedulers, autoscalers, or resource management systems
- Experience with container orchestration (Kubernetes) and sandboxed/virtualized code execution at scale
- Experience with high-performance networking, RDMA, or collective communication libraries
- Experience building observability or automated remediation for large fleets
- Experience with async Python frameworks (Trio, asyncio)
- Familiarity with reinforcement learning or large language model training workloads