SlipstreamJobsFresh Startup & VC-Backed Jobs

Research Engineer / Research Scientist, RL Frontiers

Anthropic - San Francisco, CA, United States - Hybrid - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 500,000 - 850,000 / annual

Anthropic is seeking a Research Engineer to join the RL Scaling team, which focuses on how reinforcement learning scales as models grow larger, episodes extend longer, and compute increases by orders of magnitude. This role sits at the intersection of research and engineering, requiring both algorithmic innovation and systems-level problem-solving. You will develop next-generation RL architectures and algorithms, scaling them from small-scale prototypes to frontier-scale production runs. Key responsibilities include: - Studying how RL training and sampling scale with model size, context length, and compute, identifying algorithmic and systems changes needed to maintain scaling efficiency - Developing next-generation model architectures and RL algorithms optimized for frontier-scale execution - Diagnosing why small-scale results behave differently at scale, whether due to numerical, algorithmic, or systemic factors - Building experimental infrastructure that enables fast, reproducible comparisons of architecture and algorithm variants at meaningful scale - Owning end-to-end performance of Anthropic's largest RL runs, from research code to hardware - Creating performance and cost models for proposed changes to guide scaling decisions - Investigating training dynamics at scale, including instabilities, divergence, and throughput regressions, tracing them to root cause Representative projects include characterizing how new RL algorithms scale from small to frontier models, developing and validating novel attention variants at full scale, preparing large-scale RL runs by identifying and fixing scaling bottlenecks, diagnosing loss instabilities that only appear past certain scales, and building predictive models for architecture change throughput and cost. Anthropnic operates as a single cohesive team focused on a few large-scale research efforts, emphasizing impact over incremental puzzles. The company values empirical AI research with strong communication and collaboration. REQUIREMENTS: Minimum qualifications: - Deep familiarity with modern transformer language models, including architecture, training dynamics, and large-scale optimization behavior - Hands-on experience training large models in distributed settings, understanding tradeoffs between data, tensor, and pipeline parallelism - Track record of original technical work in ML training or systems (new methods, architectures, or optimizations) demonstrated through research, open-source, or production impact - Ability to design rigorous experiments at scale with proper baselines, ablations, and statistical rigor for compute-intensive results - Quantitative reasoning about compute, memory, and communication costs of models and algorithms - Strong Python programming and proficiency in JAX or PyTorch; comfort reading and modifying code across the full stack Preferred qualifications: - Research experience in reinforcement learning, optimization, or large-scale training (published or otherwise) - Experience developing RL algorithms for language models - Experience with scaling laws or quantitative models of training efficiency - Experience designing or modifying transformer architectures beyond standard configurations - Experience scaling training to large accelerator fleets and debugging scale-specific problems - Deep understanding of numerics in large-scale training, including low-precision formats and instability sources - Familiarity with GPU/TPU performance characteristics and their impact on architecture and algorithm choices - Experience with C++ or Rust Minimum education: Bachelor's degree or equivalent combination of education, training, and professional experience in a field relevant to the role. Hybrid policy: Currently expects all staff in office at least 25% of the time; some roles may require more.

Similar roles