SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff Machine Learning Engineer

ServiceNow - Santa Clara, CA, United States - Hybrid - posted 2026-09-21

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 201,300 - 352,300 / annual

ServiceNow's AI Engineering and Delivery team is seeking a Senior Staff Machine Learning Engineer to design and build production-grade agentic AI systems embedded across the platform. You will work within the Emerging Tech group, a small senior team turning early bets on AI into strategic capabilities for customers and the company. Your core responsibilities include: **Agentic Architecture**: Design and ship multi-agent systems covering orchestration, tool use, planning loops, memory, and failure recovery that operate reliably in production environments, not prototypes. **Enterprise-Grounded Reasoning**: Build agents that leverage ServiceNow's data layer—CMDB, Workflow Data Fabric, and Knowledge Graph—to make decisions with context that frontier models lack independently. **Trust, Safety, and Governance**: Own guardrails including observability, human-in-the-loop controls, and compliance infrastructure that make autonomous systems safe to deploy at Fortune 500 scale. **Retrieval and Grounding**: Work closely with the search team to ensure agents are grounded in accurate, low-latency retrieval through RAG pipelines, hybrid search, re-ranking, and evaluation. **Model Integration and Evaluation**: Integrate frontier models (Anthropic, Google, OpenAI) into the Sense → Decide → Act → Govern architecture; evaluate trade-offs across cost, latency, and capability for production use cases. **Technical Leadership**: Set architectural patterns for the group, own hard design decisions, run design reviews, and raise the bar on agentic design and production AI discipline across engineers and principals. **Scalable Architecture Design**: Build architectures supporting large-scale deployment across hyperscalers and ServiceNow's own infrastructure, with focus on emerging model capabilities applied to real-world customer problems on short timelines. ServiceNow is building the AI control tower for business reinvention, with 85% of the Fortune 500 on its platform processing 6.5T transactions annually. Current production systems include AI Specialists autonomously resolving cases, Action Fabric for external agent integration via MCP, Project Arc with NVIDIA for governed autonomous desktop agents, and AI Control Tower with governance capabilities. **Requirements**: - 6+ years of software engineering with strong fundamentals in data structures, algorithms, and distributed systems - Hands-on depth designing, shipping, and operating agentic systems in production (multi-agent orchestration, tool calling, planning loops, memory, failure recovery)—not prototypes - Production-grade Python; systems language (Go, Java, or C++) is a plus - Working experience with frontier AI SDKs (Anthropic, Google, or OpenAI)—prompt engineering, structured outputs, and model evaluation in production settings - Familiarity with RAG and retrieval patterns in production (vector stores, hybrid search, retrieval evaluation metrics) - Track record of technical leadership: architecture ownership, code quality bar-raising, and mentoring engineers on production AI practices **Nice to Have**: - Deeper specialization in search and retrieval at scale or MLOps/model observability - Published work or open-source contributions in agentic systems or retrieval - Exposure to LLM fine-tuning or inference optimization in production

Similar roles