SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceNow is building an evaluation layer that validates AI agents before they reach customers—an automated system that scores across large volumes of agent traces. As a Senior Research Scientist in Agentic AI, you will own core parts of this platform: the pipeline that runs traces through model-based judges at scale, and the scoring logic that transforms raw output into actionable results.
Your responsibilities include:
- Design evaluation methodologies and benchmarks for agent reasoning, planning, tool use, reliability, and safety across LLM-as-a-Judge, trajectory-based, and human evaluation approaches
- Own problems end-to-end, from research question through prototype to shipped feature
- Build and harden the pipelines and scoring logic that power customer-facing evaluation systems
- Curate synthetic and real-world datasets; measure evaluator consistency and agreement with human labels
- Communicate technical findings clearly to both technical and non-technical audiences
ServiceNow operates with flexible work personas. The company is an AI control tower for business reinvention, serving 85% of the Fortune 500, and is building an AI-native culture where technology and talent work together.
QUALIFICATIONS:
- 5+ years in ML, applied AI, prompt engineering, or agentic AI, with shipped products to real users
- Strong Python skills
- Practical depth in agentic AI and context engineering: planning, reasoning, memory, tool use, retrieval, long-context
- Experience designing evaluation methodologies (not just running evaluations)
- Hands-on production work with LLM APIs—prompt engineering, structured output, cost and latency tradeoffs
- Clear communication with technical and non-technical audiences
- Good to have: Experience with AI-assisted development tools (Claude Code, Windsurf, or similar)