SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceNow is building an evaluation layer that validates AI agents before they reach customers—an automated system that scores across large volumes of agent traces. You will own core parts of this platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that turns raw output into results teams can act on.
In this role, you will design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety across LLM-as-a-Judge, trajectory-based, and human evaluation approaches. You'll take problems from research question to prototype to shipped feature, owning them end to end. You'll build and harden the pipelines and scoring logic behind customer-facing evaluation systems. You'll also curate synthetic and real-world datasets, and measure the evaluator itself for consistency and agreement with human labels.
This is a hands-on role that bridges research and production. You'll work with technical and non-technical audiences to communicate findings and drive adoption of evaluation methodologies across the organization.
REQUIREMENTS:
- 5+ years in ML, applied AI, prompt engineering, agentic AI, including shipping something real to users
- Strong Python skills
- Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context
- Experience designing evaluation methodologies, not just running evaluations
- Hands-on production work with LLM APIs—prompt engineering, structured output, cost and latency tradeoffs
- Clear communication with technical and non-technical audiences
- Good to have: Experience with AI-assisted development tools (Claude Code, Windsurf, or similar)