SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff ML Engineer, Agent Training & Environments

Labelbox - San Francisco, CA, USA - In-office - posted 2026-07-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Labelbox is building critical infrastructure for frontier AI development, operating as the RL data factory for advancing agent capabilities. The company provides three integrated solutions: an enterprise platform with annotation tools and workflow automation, a specialized data labeling service (Alignerr), and an expert marketplace connecting AI teams with skilled annotators. This Staff ML Engineer role sits at the intersection of training and infrastructure. You will design and build both the experiments and the systems that run them: RL environments where agents operate, verifiers that evaluate success, and fine-tuning pipelines that convert evaluation signals into model improvements. Key responsibilities include: - Designing and implementing RL environments for agentic tasks, including task definitions, tool surfaces, state/reset semantics, reward design, and parallel execution harnesses - Building verifiers and graders (programmatic checks, LLM judges, rubric pipelines, pass@k scoring) that make success judgments trustworthy at scale - Developing fine-tuning pipelines for SFT and RL that turn evaluation signals into measurable agent improvements - Creating eval systems that run millions of agent trajectories to measure model and product quality - Building training and serving infrastructure for frontier-scale throughput: multi-launcher orchestration, fault tolerance, cost accounting You'll need 3+ years shipping production systems, exceptional engineering throughput without quality compromise, strong system and API design judgment, and daily experience with coding agents in production. Deep Python proficiency is required, with comfort across the full stack. On the RL side, you must have fine-tuned models for agentic tasks using SFT plus at least one RL method (GRPO, PPO, DPO, etc.) in production, built environments for agents, and designed verifiers for open-ended work. The environment is startup-paced with high agency and rapid execution. You'll have clear ownership, autonomy to execute, and influence over technical direction. The role rewards those who move fast in ambiguous conditions, set technical direction, and turn prototypes into reliable systems quickly.

Similar roles