SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Mind Robotics is building robots that learn from real-world experience. The data platform team transforms raw sensor data into training signals that power robot learning. Every day, data flows in from multiple sources in different formats, volumes, and quality levels. This team owns the entire pipeline: ingesting data reliably, validating and curating it for quality and diversity, powering annotation workflows, and building automated labeling methods.
As a Software Engineer on the data platform, you'll be one of the founding engineers on the application layer team, reporting to the Head of Application Engineering. The systems work today; your role is to harden and productionize them for scale as the company ingests from a growing number of field sites and sources. This is early-stage, hands-on 0-to-1 engineering work where you'll build the pipelines and tooling that determine what data reaches the models and directly see the impact on robot behavior.
Key responsibilities:
- Build data ingestion pipelines from multiple sources, including field capture and teleop stacks
- Design and implement automatic data validation systems to catch quality issues before annotation or training
- Build and improve annotation ingestion, tooling, and workflows to increase labeling efficiency and throughput
- Own data quality and diversity metrics—build systems that track what data exists, identify gaps, and guide collection priorities
- Explore and implement automated annotation methods using computer vision and vision-language models (depth estimation, hand tracking, auto-labeling) to reduce manual labeling overhead
- Support the platform in production: debug live pipeline issues, instrument for observability, and collaborate directly with research and annotation partners
Requirements:
- 2+ years of software engineering experience building production systems
- Strong programming fundamentals and ability to work across the stack (services, data pipelines, applied ML tooling)
- Experience with at least one of: real-time or large-scale data pipelines, ML data workflows, computer vision, or data infrastructure
- Demonstrated ownership: you've taken features or systems from prototype to production and supported them in the field
- Clear communication and close collaboration with product, research, and annotation/operations partners
- Plus: hands-on experience with sensor data (video, depth, IMU, force/torque) and infrastructure to process, validate, and label at scale
- Plus: streaming or near-real-time data pipeline experience
- Plus: familiarity with ML data workflows (datasets, labeling, evaluation)
- Plus: experience running or integrating computer vision or vision-language models (depth, pose/hand tracking, open-vocabulary detection, auto-labeling)