SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Engineer, Machine Learning Systems & Reliability - Moveworks

ServiceNow - Mountain View, CA, United States - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ServiceNow (which acquired Moveworks in December 2025) is seeking a Staff Engineer to build production systems for AI-enabled product capabilities. This role sits at the intersection of ML systems, platform engineering, and site reliability engineering, partnering with ML, data, product, and infrastructure teams to create a paved path from experimentation to production. You will design and build the complete ML lifecycle in production: data and feature preparation, training, experiment tracking, evaluation, artifact and model management, serving, monitoring, feedback collection, and retraining. You'll establish continuous-delivery workflows for models, prompts, agent workflows, and data dependencies with automated quality, safety, performance, and compatibility checks. You'll implement safe rollout patterns including shadow traffic, canaries, progressive delivery, feature flags, versioned artifacts, automated rollback, and operational kill switches. A key focus is turning self-learning approaches into controlled production feedback loops—building systems for collecting outcomes, validating feedback, maintaining lineage, triggering model refreshes, comparing candidates, and promoting changes under explicit guardrails. You'll define and operate SLIs, SLOs, alerts, and error budgets across infrastructure, data pipelines, inference services, model quality, and product behavior. You'll connect model analytics and product telemetry with traditional operational signals so teams understand whether problems originate in infrastructure, data, model behavior, or the surrounding product. You'll improve scalability, availability, latency, and cost efficiency of distributed training, inference, and data-processing workloads, owning capacity planning and resource optimization including GPU resources. You'll participate in production ownership across the service lifecycle: architecture reviews, deployment, on-call, incident response, blameless postmortems, and systemic remediation. You'll build self-service platforms and automation to reduce operational toil and shorten time-to-production for ML engineers and data scientists. You'll apply LLMs or agentic automation to evaluation, troubleshooting, and operational workflows. You'll establish practical standards for cloud infrastructure, Kubernetes, infrastructure as code, observability, security, and compliance, and provide technical leadership across teams through mentoring and architecture influence without formal authority. REQUIREMENTS: - Track record of Staff-level technical ownership, typically 7+ years in software engineering, platform engineering, SRE, production engineering, or ML infrastructure - Strong software-engineering skills in Python and at least one production systems language (Go, Java, C++, or Rust) - Experience designing, operating, and troubleshooting distributed production systems, including failure analysis, capacity planning, and performance optimization - Hands-on experience with cloud infrastructure, containers and Kubernetes, infrastructure as code, CI/CD, and modern observability - Practical understanding of the ML lifecycle (training, evaluation, model deployment, serving, monitoring, versioning, retraining) and ability to collaborate with applied ML engineers or researchers - Experience distinguishing service-health problems from data-quality or model-quality problems - Familiarity with SRE practices (SLIs/SLOs, error budgets, sustainable on-call, incident management, blameless postmortems) - Strong automation and internal-customer mindset; ability to build reliable, understandable, pleasant platforms for other engineers - Excellent technical judgment and communication skills, especially navigating ambiguity and coordinating across teams during production incidents

Similar roles