SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer – SRE & AIOps

ServiceNow - Santa Clara, CA, United States - Hybrid - posted 2026-09-10

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 166,500 - 291,400 / annual

ServiceNow is seeking a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable global engineering teams to operate reliably at scale. Key responsibilities include: • Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices supporting 99.99%+ availability targets. • Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures and reduce MTTR. • Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations supporting global follow-the-sun on-call operations. • Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, developing automated runbooks that empower on-call engineers to resolve issues autonomously. • Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines enabling reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails. • Architect hybrid cloud and data center operations spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices. • Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls. • Design on-call rotation schedules, escalation policies, and incident command systems spanning different time zones, ensuring 24/7 incident response while driving post-incident review processes. • Mentor and guide junior SRE engineers and infrastructure teams on reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations. • Champion a culture of blameless incident analysis, data-driven decision-making, continuous improvement, and experimentation across engineering teams. • Reduce operational toil through systematic automation of repetitive tasks, directly improving team capacity and job satisfaction across globally distributed operations. This role combines strong hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain high reliability while minimizing operational toil. REQUIREMENTS: • 8+ years in software engineering or infrastructure operations, with 5+ years in SRE, DevOps, or cloud platform engineering roles managing large-scale distributed systems (Bachelor's degree required; or 6 years with Master's degree; or PhD with 3 years experience; or equivalent experience). • 4+ years hands-on experience designing, deploying, and operating production Kubernetes clusters at scale, including cluster design, node management, pod orchestration, resource quotas, network policies, security controls, and troubleshooting complex runtime issues. • Proven experience designing and implementing closed-loop automated remediation systems, including anomaly detection, alert correlation, runbook automation, and self-healing mechanisms that measurably reduce MTTR and on-call burden. • Extensive hands-on experience with AWS (EKS, EC2, RDS, Lambda), Azure (AKS, VMs, CosmosDB), and GCP (GKE, Compute Engine, Cloud SQL), capable of architecting multi-region solutions. Demonstrable experience across 2+ of these platforms required. • Proficiency in Infrastructure-as-Code tools (Terraform, CloudFormation, or equivalent) used to manage infrastructure at scale, and GitOps platforms. • Strong working knowledge of observability platforms, incident management systems, and log aggregation. • Solid understanding of distributed system challenges, eventual consistency, cascading failures, network partitions, and proven ability to design systems resilient to these conditions. • Experience operating or designing components of 24/7 follow-the-sun on-call models for distributed teams, including runbook development and incident response. • Hands-on experience managing both on-premises infrastructure and public cloud environments, including hybrid networking, disaster recovery, and workload migration strategies. • Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash). • Demonstrated ability to apply machine learning and AI-driven insights to infrastructure operations, including anomaly detection, predictive alerting, and intelligent remediation. Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving required. • Proven ability to drive technical decisions across teams and mentor engineers on reliability practices through credibility and technical depth. • Bachelor's degree in computer science, computer engineering, or related field (or equivalent professional experience). PREFERRED: • Kubernetes certification (CKA, CKAD, or equivalent). • Experience with service mesh platforms or advanced networking in Kubernetes environments. • Background in migrating workloads from on-premises data centers to public cloud environments. • Experience with cost optimization practices in hybrid cloud environments (reserved instances, spot instances, resource right-sizing). • Track record of mentoring infrastructure engineering teams.

Similar roles