SlipstreamJobsFresh Startup & VC-Backed Jobs

Software Engineer II, Reliability

Klaviyo - Dublin, Ireland - In-office - posted 2026-07-16

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Klaviyo is seeking a Software Engineer II, Reliability to join their Site Reliability Engineering (SRE) team in Dublin. This role focuses on ensuring Klaviyo's critical platforms are reliable, scalable, and sustainable while enabling rapid product development. You will contribute to reliability and operational excellence by working on well-scoped projects and owning services with support from senior engineers. Key responsibilities include building and operating production systems with a focus on reliability, scalability, and performance; applying software engineering principles to automate operational tasks and reduce manual toil; contributing to system design using established SRE best practices; defining and measuring SLIs and SLOs for services; improving observability through metrics, dashboards, logging, and tracing; participating in on-call rotations and responding to production incidents; assisting with incident investigation and post-incident reviews; analyzing system behavior and capacity usage; identifying reliability issues and working with teammates to address them; collaborating with product, platform, and security engineers; and writing clear operational runbooks and documentation. You are an early-to-mid career SRE comfortable operating production systems and eager to deepen your reliability engineering expertise. Required qualifications include experience operating cloud-native production systems; writing production-quality code in Python, Go, or similar languages; understanding common failure modes in distributed systems; hands-on experience with containerized workloads and Kubernetes in production; comfort with on-call rotations and diagnosing production issues; experience with observability tools; familiarity with SRE concepts like SLIs, SLOs, and error budgets; hands-on experience with infrastructure as code or declarative configuration; ability to follow incident response processes; and openness to feedback and continuous improvement. You should also have experimented with AI tools and be excited to explore new AI workflows responsibly. Nice-to-have skills include experience supporting security-sensitive systems, familiarity with AWS or other cloud providers, exposure to messaging systems like Kafka or RabbitMQ, interest in performance testing and capacity planning, and practical experience with algorithms and data structures. Klaviyo's platform is primarily built with Python and React.

Similar roles