SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 166,500 - 291,400 / annual
ServiceNow is seeking a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable global engineering teams to operate reliably at scale.
Key responsibilities include:
• Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices supporting 99.99%+ availability targets.
• Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures and reduce MTTR.
• Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations supporting global follow-the-sun on-call operations.
• Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, developing automated runbooks that empower on-call engineers to resolve issues autonomously.
• Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines enabling reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails.
• Architect hybrid cloud and data center operations spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices.
• Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls.
• Design on-call rotation schedules, escalation policies, and incident command systems spanning different time zones, ensuring 24/7 incident response while driving post-incident review processes.
• Mentor and guide junior SRE engineers and infrastructure teams on reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations.
• Champion a culture of blameless incident analysis, data-driven decision-making, continuous improvement, and experimentation across engineering teams.
• Reduce operational toil through systematic automation of repetitive tasks, directly improving team capacity and job satisfaction across globally distributed operations.
This role combines strong hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain high reliability while minimizing operational toil.
REQUIREMENTS:
• 8+ years in software engineering or infrastructure operations, with 5+ years in SRE, DevOps, or cloud platform engineering roles managing large-scale distributed systems (Bachelor's degree required; or 6 years with Master's degree; or PhD with 3 years experience; or equivalent experience).
• 4+ years hands-on experience designing, deploying, and operating production Kubernetes clusters at scale, including cluster design, node management, pod orchestration, resource quotas, network policies, security controls, and troubleshooting complex runtime issues.
• Proven experience designing and implementing closed-loop automated remediation systems, including anomaly detection, alert correlation, runbook automation, and self-healing mechanisms that measurably reduce MTTR and on-call burden.
• Extensive hands-on experience with AWS (EKS, EC2, RDS, Lambda), Azure (AKS, VMs, CosmosDB), and GCP (GKE, Compute Engine, Cloud SQL), capable of architecting multi-region solutions. Demonstrable experience across 2+ of these platforms required.
• Proficiency in Infrastructure-as-Code tools (Terraform, CloudFormation, or equivalent) used to manage infrastructure at scale, and GitOps platforms.
• Strong working knowledge of observability platforms, incident management systems, and log aggregation.
• Solid understanding of distributed system challenges, eventual consistency, cascading failures, network partitions, and proven ability to design systems resilient to these conditions.
• Experience operating or designing components of 24/7 follow-the-sun on-call models for distributed teams, including runbook development and incident response.
• Hands-on experience managing both on-premises infrastructure and public cloud environments, including hybrid networking, disaster recovery, and workload migration strategies.
• Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
• Demonstrated ability to apply machine learning and AI-driven insights to infrastructure operations, including anomaly detection, predictive alerting, and intelligent remediation. Experience leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving required.
• Proven ability to drive technical decisions across teams and mentor engineers on reliability practices through credibility and technical depth.
• Bachelor's degree in computer science, computer engineering, or related field (or equivalent professional experience).
PREFERRED:
• Kubernetes certification (CKA, CKAD, or equivalent).
• Experience with service mesh platforms or advanced networking in Kubernetes environments.
• Background in migrating workloads from on-premises data centers to public cloud environments.
• Experience with cost optimization practices in hybrid cloud environments (reserved instances, spot instances, resource right-sizing).
• Track record of mentoring infrastructure engineering teams.