SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 255,000 - 300,000 / annual
Robinhood is building an elite AI Platform & Agentic Apps team to power agent systems across the company. The mission is to democratize finance for all, and this role sits at the center of building trustworthy, scalable agentic AI in a regulated financial environment.
As a Staff Machine Learning Engineer, you will design and build the core harness that every AI agent at Robinhood runs on. This includes orchestration, tool integrations, context and memory management, and safety mechanisms. You'll work on both high-trust internal agents (used by employees) and customer-facing agents that take real action on behalf of users.
Key responsibilities:
- Design and build Robinhood's agent harness, supporting both internal and customer-facing agents with a unified platform architecture
- Ship agentic applications end-to-end, from problem definition through production deployment, and feed learnings back into the platform
- Build trajectory-level evaluation systems that score agent reasoning and decision-making, not just final answers—including tool-call correctness, planning, recovery, and multi-step task completion
- Architect action guardrails as platform primitives: least-privilege tool scoping, permission models, human-approval gates, step/budget limits, sandboxing, and rollback capabilities
- Make evals and guardrails adoptable by other teams through SDKs, CI regression gates, continuous red-teaming, and production tracing that closes the loop from real traffic back into eval sets
- Set technical standards through architecture reviews, code reviews, and mentorship; be the person who can make and defend the "don't ship" call with data
You'll be a technical anchor on a high-caliber team, collaborating with product, infrastructure, and fellow ML engineers to take ambitious ideas from zero to one and into production. This role offers rare technical depth, platform-scale impact, and the satisfaction of building systems that don't exist elsewhere.
The role is based in Menlo Park, CA with in-person attendance expected at least 3 days per week.
REQUIREMENTS:
- 10+ years of experience as a Machine Learning Engineer or ML-focused software engineer, with strong Python and distributed-systems fundamentals and a track record of shipping LLM-powered systems to production at scale
- Master's degree in Computer Science or related technical field, or equivalent professional experience
- Hands-on experience building agentic systems end-to-end (tool use, orchestration, context management, multi-step planning) on top of frontier models in production
- Deep expertise evaluating agents: trajectory-level evals, tool-call scoring, simulation environments; ability to articulate why final-answer accuracy is insufficient for systems that act
- Demonstrated expertise designing action-level guardrails (permission/tool-scoping models, approval gates, blast-radius controls, sandboxing) for agents in systems where mistakes have consequences
- Rigor in evaluation methodology: golden datasets, rubric and LLM-as-judge grading and their failure modes, statistical significance with small N, offline-to-online metric correlation, eval data versioning and contamination control
- Proven ability to build platforms, not just models: shipped eval, safety, or agent tooling that other engineering teams adopted; judgment to know when to build versus buy