SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 500,000 - 850,000 / annual
Anthropic is seeking a Research Engineer to join the Reinforcement Learning Engineering team. You will build and maintain the critical algorithms and infrastructure that enable researchers to train advanced AI models like Claude using RLHF and related techniques.
In this role, you'll focus on improving the performance, robustness, and usability of training systems that directly support breakthrough research in AI capabilities and safety. Your responsibilities include profiling and optimizing the RL pipeline, building systems to detect training issues early, adapting finetuning systems for new model architectures, eliminating performance bottlenecks (e.g., Python GIL contention), diagnosing and fixing training slowdowns, and implementing stable, production-ready versions of novel training algorithms proposed by researchers.
You'll work closely with a team of researchers and engineers who are passionate about building beneficial AI systems. The role emphasizes systems thinking, collaborative problem-solving, and a bias toward impact. You'll be energized by the challenge of empowering research teams to move faster and iterate more effectively.
Minimum qualifications include 4+ years of software engineering experience and a bachelor's degree in a relevant field. Strong candidates often have experience with high-performance distributed systems, large-scale LLM training, Python, and implementing finetuning algorithms. The company expects hybrid work with at least 25% office time at one of the three listed locations, though some roles may require more in-office presence. Visa sponsorship is available.