SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, Code RL

Anthropic - San Francisco, CA, United States - Hybrid - posted 2026-07-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Anthropic is seeking a Staff Software Engineer to lead reinforcement learning infrastructure efforts that power Claude's coding capabilities. This role sits at the intersection of research and engineering, with significant latitude to set technical direction and establish standards for a growing team. You'll embed with research teams on a rotational basis, understanding their engineering needs and designing the frameworks, APIs, and infrastructure that accelerate their work. Your responsibilities include designing widely-adopted APIs and abstractions that researchers and engineers build upon, working directly in research codebases to improve reliability and structure without slowing progress, and anticipating silent failure modes through principled system design, type safety, and testing. The technical scope spans client-side sandboxed execution for agentic RL environments, large-scale data processing pipelines, production dataset lifecycle management, and the frameworks researchers use to build environments. You won't own all of this alone—you'll focus on areas where your depth matters most, rotating off after establishing systems that teams can maintain independently. Key responsibilities include designing intuitive, safe APIs with careful attention to interface legibility; embedding with research teams to understand needs and transfer ownership; improving reliability and structure in research codebases; preventing failure modes structurally; contributing to production RL system reliability including monitoring and triage tooling; and helping define engineering standards, review practices, and design patterns for the team. You should have deep Python expertise including static typing, async/concurrency patterns, and performance optimization. A track record of designing APIs or frameworks that teams adopted is essential. You need demonstrated ability to work productively in large, evolving research codebases you didn't write, anticipate failure modes, and communicate system designs clearly to collaborators with varied backgrounds. Comfort with ambiguity and ability to scope work from loosely defined problems is critical. Preferred experience includes building ML research infrastructure, familiarity with RL concepts and LLM training pipelines, large-scale distributed systems, client libraries for sandboxed execution, dataset lifecycle management, extensible plugin systems, embedding with other teams, defining code standards across teams, prior technical leadership, or maintaining open source projects.

Similar roles