SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Anthropic's RL Data Platform team builds the infrastructure that powers human feedback collection and training data pipelines for Claude. This is a full-stack, ownership-heavy role on a small, senior engineering team.
You'll design and ship web interfaces used by thousands of expert annotators, build the backend services and data pipelines that route model samples to humans and return structured feedback to training, and work directly with RL researchers to understand their data needs. The role spans the entire stack: TypeScript/React frontends for annotation interfaces, Python backends and data pipelines, and the operational systems that keep everything running reliably against live model endpoints.
Key responsibilities include:
- Designing and building feedback collection interfaces for human annotators, domain experts, and internal researchers
- Building and maintaining backend services, APIs, and data pipelines that handle the flow from model samples to structured training feedback
- Owning reliability, latency, and usability of systems running continuously against live endpoints
- Partnering with RL researchers to translate loosely specified data needs into well-scoped collection campaigns and tooling
- Building dashboards, monitoring, and inspection tools so researchers can self-serve on data quality and throughput
- Identifying and removing bottlenecks in the pipeline from "we want this data" to "it's in the training mix"
You'll scope your own projects, make architectural decisions, and see them through to production. The team values engineers who treat researchers as users, prioritize reliability, and care deeply about data quality.
Required: strong full-stack skills (TypeScript/React frontend, Python backend), experience designing and operating backend services and data pipelines that other teams depend on, track record of owning projects end-to-end from ambiguous brief to production, comfort working with technical stakeholders whose needs evolve, effective use of AI tools, and care about societal impact.
Preferred: experience with annotation/labeling/evaluation tooling, RLHF or human-feedback pipelines, shipping researcher-facing internal tools, running experiments on data collection interfaces, working with crowdworker platforms at scale, or familiarity with LLM training and evaluation.