SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Salary: USD 98,000 - 148,500 / annual
PagerDuty is a leader in Digital Operations Management, trusted by over 13,000 organizations (including 60 Fortune 100 companies) to deliver real-time incident detection, response, and resolution. The company operates a global platform that processes millions of events and alerts daily, enabling teams across development, IT, security, and customer service to keep their digital operations running smoothly.
As a Site Reliability Engineer I on the Core Infrastructure team in Atlanta, you will help build and operate the foundational infrastructure powering PagerDuty's real-time platform. You'll work at the intersection of platform evolution and operational excellence, focusing on network, compute, and ingress infrastructure while scaling and hardening existing systems.
Key responsibilities include:
- Supporting and improving foundational infrastructure including networking, compute platforms, Kubernetes clusters, and ingress/traffic management systems
- Contributing to reliability and scalability by hardening existing systems and supporting rollout of new infrastructure capabilities
- Participating in agile rituals (standups, planning, retros) and communicating progress and risks early
- Staying current on technical trends to suggest innovative tools and approaches
- Monitoring system health using metrics, logs, and alerts
- Participating in 24/7 on-call rotations to detect, respond to, and resolve incidents
Required qualifications include 0-1+ years of SRE, DevOps, or Platform Engineering experience; hands-on Linux production operations; networking fundamentals (load balancing, DNS, TLS, ingress); container orchestration (Kubernetes/EKS); cloud-native infrastructure experience (AWS/GCP/Azure); proficiency in at least one programming language (Python, Ruby, Go); and Infrastructure as Code experience (Terraform, CloudFormation).
Preferred qualifications include AWS networking concepts (VPCs, subnets, routing, security groups), production Kubernetes platform experience, monitoring/observability platforms (DataDog, New Relic, Prometheus, Grafana), and familiarity with service meshes and ingress controllers (Envoy, Istio, NGINX).
PagerDuty operates a hybrid model with offices in 8 major cities. The company values potential and encourages applications even if you don't meet every requirement.