SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 250,000 - 450,000 / annual
Roboflow is building the visual cortex for AGI, with millions of developers and half of the Fortune 100 using the platform for computer vision applications ranging from cancer research to construction safety to drone guidance. The company has raised over $100 million from Google Ventures, Craft Ventures, Sam Altman, and Y Combinator.
Roboflow Labs is the company's applied research unit, focused on training open models (RF-DETR), building benchmarks grounded in real enterprise use cases (RF100-VL), and partnering with frontier labs to deliver training data that advances AI capabilities.
As a Research Scientist at Roboflow Labs, you will sit at the intersection of benchmarks, models, and frontier lab researchers. Your core responsibilities include:
- Design datasets, evaluations, and RL environments targeting measurable gaps in frontier vision-language, multi-modal, robotics, and coding models
- Run structured failure analyses on partner labs' models against Roboflow's real-world task distribution and write remediation specs detailing what data will close each gap
- Own the research behind new benchmarks that the team publishes
- Validate data quality empirically through fine-tuning, ablation studies, and measurement to ensure every dataset shipped comes with evidence of effectiveness
- Contribute to Roboflow's open models (RF-DETR and successors) and represent Labs within the research community
- Work directly with research scientists at frontier labs as a technical peer, translating their model roadmaps into concrete data programs
You will work in-person in San Francisco. Roboflow is a distributed company with hubs in San Francisco, New York City, Singapore, and Brazil, hosting full company retreats twice per year.
REQUIREMENTS:
- PhD or equivalent research experience in computer vision, multi-modal learning, machine learning, or closely related field, with a track record of published work or shipped research
- Hands-on experience with post-training: fine-tuning, RL, evaluation design, or data curation for large models
- Strong engineering fundamentals — you write code for your own experiments and can hand it to an engineer for productionization
- Clear written communication; ability to explain failure modes to lab researchers and data specs to annotators
- Nice to have: experience at a frontier lab or as a data/evaluation partner to one; experience with vision-language models, robotics learning, or agentic coding evaluations