SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer II

PagerDuty - Atlanta, GA, United States - Hybrid - posted 2026-09-01

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 113,000 - 171,600 / annual

PagerDuty is a leader in Digital Operations Management, trusted by over 13,000 organizations including 60 Fortune 100 companies to deliver digital experiences and manage incidents in real time. The company operates a hybrid model with offices in 8 major cities and is expanding its platform using AI/ML and automation. As a Site Reliability Engineer II on the Core Infrastructure team in Atlanta, you will help build and operate the foundational infrastructure powering PagerDuty's real-time digital operations platform. The systems you support handle millions of events and alerts daily, enabling customers to detect, respond to, and resolve incidents quickly and reliably. You will work at the intersection of platform evolution and operational excellence, building and evolving foundational network, compute, and ingress infrastructure while scaling and hardening existing systems. Your work will directly impact the reliability, scalability, and security of services customers depend on as PagerDuty grows across products, regions, and use cases. Key responsibilities include supporting and improving foundational infrastructure (networking, compute platforms, Kubernetes clusters, ingress/traffic management); contributing to core platform reliability and scalability by hardening systems and supporting new infrastructure rollouts; participating in agile rituals and communicating progress/risks early; staying current on technical trends to suggest innovative tools and approaches; and monitoring system health using metrics, logs, and alerts while participating in 24/7 on-call rotations. Required qualifications: 3+ years in Site Reliability Engineering, DevOps, or Platform Engineering; hands-on experience operating Linux-based systems in production; working knowledge of networking fundamentals (load balancing, DNS, TLS, ingress traffic flow); experience with container orchestration (EKS, Kubernetes); cloud-native infrastructure experience (AWS, GCP, Azure); proficiency in at least one programming language (Python, Ruby, Go); and Infrastructure as Code experience (Terraform, CloudFormation). Preferred qualifications include AWS cloud networking concepts (VPCs, subnets, routing, security groups, load balancers); production Kubernetes platform experience; monitoring/observability/logging platforms (DataDog, New Relic, SumoLogic, Splunk, Prometheus, Grafana); and familiarity with service meshes, ingress controllers, or API gateways (Envoy, Istio, NGINX).

Similar roles