SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Intelligence is a product lab company behind DesignArena, a benchmark platform with 5.2M+ users that evaluates frontier AI model capabilities across design, web development, game development, image, video, audio, and slide generation. The company is backed by Tier 1 VCs including Index Ventures, Y Combinator, and SV Angel, with a talent-dense team (11 from Harvard and Berkeley). Intelligence is trusted by frontier model providers like OpenAI to rigorously evaluate state-of-the-art multimodal models.
In this role, you will invent new ways to detect and scale the data that makes frontier AI models smarter. As frontier models continue to improve, model capability is increasingly driven by the quality of training data. You will develop scalable systems that generate, curate, and improve high-quality supervision across coding, multimodal, and agentic tasks. You have an opportunity to be forward-deployed and work directly with researchers at frontier labs to creatively scale new model capability strategies.
Key responsibilities include: training and improving preference, reward, and ranking models from millions of human interactions; developing infrastructure for large-scale experimentation, model training, and specializing in online evaluation techniques; and designing systems that transform human preference data into reliable signals for downstream model evaluation and training.
The role is based at Levi's Plaza in San Francisco. The work schedule is Sunday-Friday, with Saturdays off. The company sponsors visas and handles relocation.
Requirements: Strong STEM background in Computer Science, Machine Learning, Data Science, Statistics, Math, Engineering, Physics, or a related field. Experience building ML systems, including training preference models, building data pipelines, or working on ML infrastructure in production. Deep understanding that data quality is the bottleneck to superintelligence. Interest in how AI learns from human feedback and solving problems at the intersection of human-model interaction.