SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer, RL Environments

Scale - San Francisco, CA, United States - Hybrid - posted 2026-09-02

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Scale AI is hiring a Staff Software Engineer to own the technical foundation for reinforcement learning (RL) environments at scale. This is a hands-on role where you'll design and build the platform infrastructure that enables Scale to create, run, verify, and deliver thousands of RL environments reproducibly and cost-effectively. You'll work on both platform architecture and deep technical implementation. On the platform side, you'll design sandboxed execution systems, environment packaging and versioning, rollout orchestration, trajectory capture, verifier frameworks, and authoring surfaces that let engineers and domain experts build environments without reinventing infrastructure. On the environments side, you'll instrument real applications, design task suites that expose specific capability gaps, and build graders that hold up under adversarial optimization. RL environments are the center of gravity for frontier AI training: the difference between a model that demos well and one that reliably completes long-horizon work is almost always the quality of the environments and reward signals it was trained against. An RL environment is a full-stack engineering problem—a containerized world with real dependencies, real state, real tools, and a grader that must be correct even when agents creatively break it. Building thousands of them reproducibly, cheaply, with trustworthy reward signals and throughput measured in millions of rollouts is a systems problem very few have solved. You'll set technical direction across multiple teams while remaining hands-on—you'll be the person who writes the hard parts. This requires deep expertise in distributed systems, containerization (Docker, Kubernetes, gVisor, Firecracker), high-throughput backend systems (orchestration, job scheduling, queuing, data pipelines), and hands-on experience building with LLMs (agent loops, tool calling, eval harnesses). You'll need strong Python skills, comfort in at least one other part of the stack (TypeScript/React, Go, Rust), and 8+ years of software engineering experience with strong fundamentals in system design, data structures, and algorithms. Preferred experience includes building RL environments, agentic benchmarks, or eval harnesses; familiarity with post-training methods (RLHF, RLAIF, RLVR, PPO-family algorithms); experience designing verifiable reward signals and defending against reward hacking; and prior technical leadership at staff level in fast-moving, ambiguous environments.

Similar roles