SlipstreamJobsFresh Startup & VC-Backed Jobs

RL Environments Engineer

Bespoke Labs - Mountain View, CA, United States - In-office - posted 2026-08-27

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 250,000 - 300,000 / annual

Bespoke Labs is an applied AI research lab focused on data and RL environment curation for training and evaluating agents. The company has recently curated Open Thoughts, one of the best open reasoning datasets used by frontier labs, and trained specialized models like Bespoke-MiniChart-7B and Bespoke-MiniCheck. This is a delivery-focused role. You will build the machinery that turns environment ideas into hundreds or thousands of validated agentic coding tasks. Rather than studying environments in the abstract, you will own the systems that mass-produce them, design complex coding worlds for agents to train inside, and continuously push throughput. Success is measured on the volume and quality of environments shipped. Key responsibilities: - Build environment-generation pipelines that produce RL environments programmatically, including templating, automated grading, verification, and QA. - Create high-fidelity coding environments around real codebases, with authentic conventions, dependencies, tooling, and technical debt. - Scale agentic task creation from handfuls to hundreds and thousands of validated tasks through automation. - Build internal tools and infrastructure that remove bottlenecks in environment production. - Own the full task lifecycle: prompt design, environment setup, grader logic, running frontier models, failure analysis, and iteration. - Defend quality at scale by catching reward hacking and grader loopholes, establishing verification standards. - Direct frontier coding agents to build and validate environments faster, judging output and catching subtle failures. Requirements: - Demonstrated record of shipped volume: you have built agentic coding tasks or environments and can show how many you personally drove and their production cost. - Experience scaling output through automation rather than manual labor. - Strong software engineering fundamentals and fluency in multiple production languages. - Real production software experience: large codebases, build systems, testing, deployment, on-call, root cause analysis. - Adversarial mindset: you anticipate how models would cheat graders and fix those loopholes. - Clear understanding of frontier coding agents' capabilities and limitations. - Ownership mentality: you build, debug, and ship with minimal supervision. Bonus qualifications: - Experience with RL training systems, post-training, verifiers, or tool-use harnesses. - Background in developer tooling, CI/CD, sandboxes, or code execution infrastructure. - Large-scale automated test generation, fuzzing harnesses, or benchmark suite experience. - Contributions to public agentic benchmarks like Terminal-Bench. - Open-source work that others depend on.

Similar roles