SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Rohlik is Central Europe's leading e-grocer, operating across five countries with over a million customers. The company is building autonomous fulfillment capabilities through AI, robotics, and computer vision. This role sits on a newly created robot-learning team focused on capturing and labeling human manipulation data to train robot policies for warehouse operations.
You will own four core deliverables:
1. **Capture Specification**: Design the camera rigs, mounts, calibration protocols, and quality gates that determine which warehouse footage is trainable. You write the spec before hardware is purchased and enforce it on the floor.
2. **Ground Truth Annotation**: Hand pose is the critical signal for transferring human demonstrations to robot policies. You will own how pose is measured and labeled, accounting for real-world challenges like work gloves, occlusion, and cold warehouse conditions where published models typically fail.
3. **Evaluation Harness**: Build the system that determines which hours of footage are suitable for training. You set and maintain the trainability bar as data volume grows.
4. **First Training Runs**: Post-train open-source robot-learning models on Rohlik's proprietary dataset and run initial task evaluations.
You will work alongside an operations lead who manages floor capture and a data engineer who owns the pipeline. This is not a research role—the focus is making a corpus trainable and training the first policies on it. You are expected to stay current with the rapidly evolving robot-learning field, reproduce published claims, and steer data capture decisions before budget is spent.
The team uses AI agents (Claude Code, Devin) for implementation work; you direct these tools for harness code, data plumbing, and reproduction scripts while maintaining quality standards.
**Requirements:**
- Strong PyTorch and computer vision expertise, particularly 3D geometry, camera calibration, pose estimation, and SLAM fundamentals. Must have hands-on experience debugging extrinsic calibration drift.
- Ability to read and reproduce published papers. The pipeline is built from published work; you validate which claims survive real warehouse conditions.
- Deep familiarity with the robot-learning stack and policy classes. You must understand what training pipelines demand from data before data is collected.
- Practical dataset experience: you have trained on data you collected yourself and understand failure modes that emerge at training time.
- Comfort working as the sole ML voice on the team. You make calls, document them, and revisit based on pilot data.
- Particularly relevant: hand-pose estimation or egocentric video in production systems; multi-camera rig synchronization; experience post-training robot-learning models; ROS 2 familiarity.