SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceNow is seeking a Senior Research Engineer/Scientist to lead research and development of agentic AI systems. This role sits at the intersection of LLM training, agentic architectures, and AI evaluation—core to ServiceNow's AI control tower platform.
You will own the full research lifecycle: LLM mid-training and post-training research (continued pretraining, SFT, preference optimization, RL), with responsibility for data-mixture decisions and understanding how training changes affect downstream agent behavior. You'll research and prototype novel agentic architectures and algorithms spanning planning, reasoning, memory, skills, tool use, retrieval, and multi-agent collaboration, pushing beyond existing approaches.
A key part of the role is designing research harnesses and experimentation methodologies that enable systematic experimentation, trajectory analysis, reproducibility, and rigorous model/checkpoint/architecture comparison. You'll define evaluation methodologies and develop novel benchmarks for agent reasoning, planning, tool use, reliability, factuality, and safety—establishing rigorous approaches for LLM-as-a-Judge, trajectory-based, and human evaluation.
You'll identify systematic model and agent failure patterns, determine root causes, and translate insights into new research directions or improvements in training data, architecture, context, or evaluation. Critically, you'll independently identify research problems, formulate novel hypotheses, and drive projects from research idea to validated prototype and measurable impact. You'll collaborate with engineering to transition successful approaches into production and communicate results through publications, patents, or open-source work.
This is a hands-on research role requiring deep technical ownership and the ability to operate independently while influencing cross-functional teams.
REQUIREMENTS:
- 5+ years of experience in machine learning, deep learning, AI research, or related field, with demonstrated applied research experience and track record of independently driving research projects.
- Hands-on experience with LLM training and post-training, including one or more of: continued pretraining, SFT, preference optimization, or RL; experience making training-data decisions and understanding training dynamics and failure modes at scale.
- Strong Python and advanced PyTorch expertise, with experience modifying models, training pipelines, or research infrastructure to support novel experimentation.
- Strong practical depth in agentic AI and context engineering, including planning, reasoning, memory, skills, tool use, retrieval, long-context processing, and knowledge grounding.
- Experience designing evaluation methodologies (not just running evaluations), including benchmark design, trajectory-based evaluation, LLM-as-a-Judge, human evaluation, and ability to determine which metrics and methodologies are appropriate for a research question.
- Demonstrated research ownership and impact, with ability to identify important research questions, develop novel hypotheses, conduct rigorous experiments, and communicate findings and implications effectively to technical researchers, engineers, and executive stakeholders.