SlipstreamJobsFresh Startup & VC-Backed Jobs

Site Reliability Engineer 2

Kong - Washington, DC, United States - In-office - posted 2026-08-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Kong is seeking a Site Reliability Engineer 2 to join its global Platform SRE team, responsible for building, operating, and scaling Kong's multi-region SaaS platform that powers API connectivity worldwide. This is a hands-on role focused on running production systems at scale, automating operations, and continuously improving performance and resilience. Key responsibilities include operating and scaling Kong's global SaaS platform (Konnect) across AWS, GCP, and Azure, ensuring reliability, availability, and performance across regions. You will build and maintain Kubernetes-based infrastructure using Terraform/Terragrunt, Helm, and ArgoCD, and design multi-region data and caching layers including PostgreSQL, Redis, ClickHouse, and Druid for high availability and low latency. You'll develop and maintain CI/CD pipelines and GitOps workflows to automate service delivery, enhance observability through Datadog, Prometheus, Grafana, and Thanos while defining and tracking SLOs, and collaborate with development and security teams to ensure smooth SaaS operations in compliance with reliability and regulatory standards. The role includes participation in a global 24/7 on-call rotation and driving continuous improvement of operational playbooks and postmortem practices. Required qualifications include a BS in Computer Science or equivalent practical experience, proven experience managing SaaS or PaaS systems at enterprise scale, deep expertise in Kubernetes including debugging and fault tolerance design, strong proficiency with Infrastructure as Code tools like Terraform, and experience with CI/CD pipelines and GitOps workflows. You should have proficiency in Go, Python, or Bash for automation, solid understanding of Linux/Unix systems and networking, and experience with API gateway technologies. Bonus qualifications include hands-on experience with Kong Gateway or Kong Mesh, experience operating ClickHouse or Druid, managing PostgreSQL and Redis in multi-region configurations, AWS networking knowledge, and understanding of disaster recovery and compliance-driven reliability practices.

Similar roles