SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 150,000 - 200,000 / annual
Contra Labs is a human-centered research lab studying how creatives use AI and how tools can better serve creative work. The company connects a global network of top creative talent with frontier AI labs and product companies, grounding research in real creative workflows to help define quality standards across design, video, imagery, and beyond. This work powers benchmarking, evaluation, and post-training for leading models and applications.
You will own project delivery and operations across multiple concurrent client engagements: post-training dataset creation, benchmarking programs, and ongoing evaluation partnerships. Each project spans different AI labs, creative-tool companies, and research teams with varying requirements, timelines, and quality standards.
You will operate as the connector between three groups: the client needing the data, the creative experts producing it, and the internal engineering team building tooling. Your responsibilities include:
**Methodology**: Design annotation protocols, rubrics, and evaluation frameworks for creative and multimodal outputs.
**Delivery**: Manage scope, evaluator staffing, quality routing, timelines, and client relationships from kickoff through final readout.
**Quality**: Ensure data quality and inter-rater reliability; hold the bar and execute real-time course corrections when needed.
**Analysis & Storytelling**: Conduct qualitative and quantitative analysis, translating findings into clear recommendations, reports, and workshops that drive client action.
**Systems & Process**: Turn repeated work into clean, measurable processes and partner with engineering on tooling that accelerates and improves reliability.
**Market Narrative**: Develop case studies and benchmark publications that clarify Contra Labs' market positioning.
You will work across several projects simultaneously, requiring high ownership, low ego, and bias toward action. You ship v1, learn rapidly, and do what it takes to deliver.
**Requirements**
- 3+ years designing studies, running evaluations, analyzing data, and presenting findings
- Experience with data labeling, annotation, human-in-the-loop systems, or AI/ML evaluation workflows
- Track record managing multiple concurrent projects involving annotation, evaluation, or QA work
- Demonstrated point of view on creative quality and what "good" looks like
- Must be based in San Francisco, California
**Bonus qualifications**
- Strong research background in both qualitative and quantitative methods
- Master's or PhD in HCI, cognitive science, psychology, design research, or related field (or equivalent industry depth)
- Statistical methods for evaluation: inter-rater reliability, sampling, regression, significance testing
- Experience with RLHF, LLM evaluation, or benchmark design
- Background in ethnography, UX research, or design-study methods
- Startup or 0-to-1 experience