SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Intelligence is a product lab company building evaluation platforms and data engines for frontier AI models. The company operates DesignArena (5.2M+ users in 8 months), a benchmark for AI-generated visuals referenced by Andrew Ng, Elon Musk, and Demis Hassabis. Intelligence is trusted by leading frontier model providers like OpenAI to rigorously evaluate state-of-the-art multimodal models across design, web dev, game dev, image, video, audio, and slide generation. The team is talent-dense (11 from Harvard + Berkeley) and backed by Tier 1 investors including Index Ventures, Y Combinator, SV Angel, and Figma Ventures.
In this role, you will build the machine learning systems powering the evaluation platform and next-generation data engine. You'll train human preference models, build large-scale data pipelines, and develop infrastructure that transforms millions of human interactions into reliable signals for evaluating and improving frontier AI models. You have the opportunity to be forward-deployed and work directly with researchers at frontier labs to creatively scale new model capability strategies.
Key responsibilities:
- Train and improve preference, reward, and ranking models from millions of human interactions
- Develop infrastructure for large-scale experimentation, model training, and specialize in online evaluation techniques
- Design systems that transform human preference data into reliable signals for downstream model evaluation and training
The role is based at Levi's Plaza in San Francisco with a Sunday-Friday work schedule (Saturdays off). The company sponsors visas and handles relocation.
Requirements:
- Strong STEM background: degree in Computer Science, Data Science, Statistics, Math, Engineering, Physics, or related field
- Experience building ML systems in production: trained preference models, built data pipelines, or worked on ML infrastructure
- Interest in how AI learns from human feedback and solving problems at the intersection of human-model interaction