SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 158,400 - 210,000 / annual
Niantic Spatial is building the future of physical AI, powered by a proprietary database of over 30 billion posed images. The company's mapping technology enables spatial intelligence for robotics, public sector, and energy/industrial markets.
The Real-World Test Lab is a new team that bridges the gap between benchmarks and real-world field conditions. Rather than relying on leaderboards, the lab brings customer environments, devices, and hardest conditions into Niantic's walls to validate that every release meets production-grade standards. The lab serves as Niantic's first and most demanding customer, pushing reconstruction, localization, and spatial understanding to their limits.
As AI Automation Engineer, you will report to the Director of the Real-World Test Lab and own the automation system that produces evidence for the entire company. Today evaluations are manual, inconsistent, and slow; you will design and build the end-to-end automation that runs them on every relevant release with no human intervention, turning one-off experiments into an always-on service.
Key responsibilities:
- Build the Evaluation Machine: Own automation for environment setup, run orchestration, artifact capture, and result collection. Make reruns free to enable constant testing.
- Make Results Comparable: Instrument scorecards so results can be compared across product versions, devices, and capture conditions, preserving data lineage.
- Automate Agent-Driven Workflows: Build agent workflows that exercise priority customer use cases at realistic scale across difficulty spectrums, with honest assessment of where agents cannot yet replace human judgment.
- Kill Manual Work: Convert one-off experiments into standing protocols that run on every relevant release; automate recurring inspection and route genuinely manual work to scalable human review.
- Make Failures Actionable: Produce diagnostics precise enough that findings reach owners with reproducible cases and attached data.
- Prevent Data Bottlenecks: Work with the AI Data Manager to ensure every run is reproducible from known dataset state without manual data handling.
Success milestones: 30 days—one priority workflow evaluated end-to-end with comparable scorecard; 60 days—evaluations trigger automatically on releases with diagnostics routing failures to owners; 90 days—three priority workflows under standing automated evaluation with version-over-version comparison.
You believe evaluation is engineering, not process. You've built systems that test other systems and understand the difference between a script that works locally and infrastructure a team can trust. You are intellectually honest, designing tests that produce trustworthy answers and refusing to let benchmarks imply more than evidence supports. You are pragmatic, finding the fastest path to defensible results without one-way doors. You reach for agents and LLM-driven automation before headcount, and you're rigorous about verification. You build for other people—your tooling is documented, failures are legible, and colleagues use what you build independently.
REQUIREMENTS:
- Built and maintained production-grade automation or test infrastructure that other engineers relied on daily
- Experience evaluating systems where correctness is graded rather than binary (quality, accuracy, or latency thresholds rather than pass/fail)
- Strong Python with fluency in orchestration, CI-style pipelines, cloud storage, and reproducible environments
- Worked with backend services and APIs you did not own, integrating without becoming a bottleneck
- Written technical findings clearly enough that non-authors could act without meetings
- Bachelor's degree in a relevant field or equivalent experience
NICE TO HAVE:
- Built agent-based or LLM-driven automation for tasks previously requiring human judgment
- Worked on evaluation or benchmarking for computer vision, 3D reconstruction, or spatial systems
- Instrumented dashboards or scorecards used by leadership for release decisions
- Operated data provisioning or registry infrastructure across multiple environments and access models