SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Rhoda AI is building the next generation of generalist intelligent robots with a full robotics stack spanning high-performance hardware, robot systems, infrastructure, and state-of-the-art foundation world models. The company has raised over $450M and is investing aggressively in model research, infrastructure, hardware development, and manufacturing scale-up.
In this role, you will own end-to-end robot model evaluation for the Research team. Your mission is to turn research questions into consistent, high-quality, repeatable evaluations that enable fast iteration and trusted results.
Key responsibilities include:
- Translating research intent into clear evaluation protocols, trial plans, and success criteria
- Owning execution end-to-end: model handoff, station readiness, pilot execution, QA, and results delivery
- Training and managing evaluation pilots to ensure consistency across people, shifts, and stations
- Distinguishing model failures from hardware, setup, operator, or data-quality issues
- Maintaining evaluation setups, resets, randomization, metadata, and experiment traceability
- Tracking quality, throughput, and bottlenecks; continuously improving the evaluation process
- Partnering closely with Research, Robot Data, and Evaluation Platform teams
Success means researchers can hand off a model and research question and receive a trusted, standardized evaluation result with sufficient trials and QA.
Requirements:
- Some understanding of robotics and/or ML experimentation
- Computer science background or hands-on experience with coding
- Rigorous, detail-oriented mindset with ability to understand the intent behind an experiment, not just execute instructions
- Strong hands-on execution and ownership capabilities
- Experience in robotics testing, data collection, lab operations, or QA (preferred)