SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Software Engineer - SRE & AIOps

ServiceNow - Atlanta, GA, United States - Hybrid - posted 2026-09-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

ServiceNow is seeking a Staff Software Engineer - SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable global engineering teams to operate reliably at scale. You will combine hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership capabilities. Key responsibilities include: - Deploy, operate, and troubleshoot production Kubernetes clusters across hybrid and multi-cloud environments, maintaining operational standards and supporting high-velocity application deployments. - Implement and maintain closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures, leveraging automation frameworks and machine learning insights to reduce MTTR and on-call burden. - Contribute to the design and evolution of SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations that support global on-call operations. - Develop and maintain SLO frameworks, alerting policies, and automated runbooks that empower on-call engineers to resolve issues autonomously while managing alert fatigue. - Build and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible infrastructure deployments across hybrid and multi-cloud environments with security and compliance guardrails. - Support hybrid cloud and data center operations, including on-premises infrastructure, public cloud environments, and workload optimization across multi-region deployments. - Contribute to adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices and network security controls. - Support on-call rotation operations and incident response processes across different time zones, helping develop runbooks and contributing to post-incident reviews. - Mentor junior SRE engineers on reliability patterns, incident investigation techniques, and automation best practices. - Champion a culture of blameless incident analysis, data-driven decision-making, and continuous improvement through knowledge sharing and documentation. - Identify and systematically automate repetitive operational tasks, from infrastructure provisioning to incident response, improving team efficiency and capacity. REQUIREMENTS: - 8+ years in software engineering or infrastructure operations, with 3+ years in SRE, DevOps, or cloud platform engineering roles (or 6 years with Master's degree, or PhD with 3 years experience, or equivalent work experience). - Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience). - 2+ years of hands-on experience working with production Kubernetes clusters. - Solid hands-on experience operating production Kubernetes clusters, including deployment models, pod orchestration, resource management, network policies, and troubleshooting runtime issues. - Demonstrated experience designing and implementing automated remediation systems, including alert automation, runbook development, and self-healing mechanisms. - Strong hands-on experience with AWS (EKS, EC2, RDS) and/or Azure (AKS, VMs) or GCP (GKE), with understanding of core SRE-related services. - Proficiency in at least one Infrastructure-as-Code tool: Terraform, CloudFormation, or equivalent. - Working knowledge of observability platforms, incident management systems, and log aggregation tools. - Understanding of distributed system challenges, fault tolerance, and resilience patterns. - Experience participating in on-call rotations and understanding 24/7 operational models, runbook development, and escalation procedures. - Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, or Bash). - Ability to work effectively with infrastructure and application teams, contribute to technical discussions, and help drive reliability improvements. - Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. - Demonstrated commitment to reliability engineering and continuous improvement through hands-on contributions. PREFERRED: - Kubernetes certification (CKA, CKAD, or equivalent). - Experience with service mesh technologies or advanced Kubernetes networking. - Background in cloud migration or infrastructure modernization projects. - Experience with cost optimization in cloud environments. - Track record of implementing automation solutions that significantly reduced operational toil.

Similar roles