SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
ServiceNow is seeking a Staff Software Engineer - SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable global engineering teams to operate reliably at scale.
You will combine hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership capabilities. Key responsibilities include:
- Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments.
- Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden.
- Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations.
- Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue.
- Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails.
- Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments.
- Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls.
- Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews.
- Mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices.
- Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation.
- Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity.
REQUIREMENTS:
- 8+ years in software engineering or infrastructure operations, with 3+ years in SRE, DevOps, or cloud platform engineering roles (or 6 years with Master's degree, or PhD with 3 years experience, or equivalent work experience).
- Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience).
- 2+ years of hands-on experience working with production Kubernetes clusters.
- Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues.
- Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self-healing mechanisms.
- Strong hands-on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE-related services.
- Proficiency in at least one Infrastructure-as-Code tool: Terraform, CloudFormation, or equivalent.
- Working knowledge of observability platforms, incident management systems, and log aggregation tools.
- Understanding of distributed system challenges, fault tolerance, and resilience patterns.
- Experience participating in on-call rotations and understanding 24/7 operational models, runbook development, and escalation procedures.
- Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, or Bash).
- Ability to work effectively with infrastructure and application teams, contribute to technical discussions, and help drive reliability improvements.
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving.
- Demonstrated commitment to reliability engineering and continuous improvement through hands-on contributions.
PREFERRED:
- Kubernetes certification (CKA, CKAD, or equivalent).
- Experience with service mesh technologies or advanced Kubernetes networking.
- Background in cloud migration or infrastructure modernization projects.
- Experience with cost optimization in cloud environments.
- Track record of implementing automation solutions that significantly reduced operational toil.