SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Klaviyo is seeking a Lead Software Engineer to provide technical leadership for the Messaging Infrastructure organization, which powers billions of real-time, low-latency messages across multiple channels (email, SMS, push, WhatsApp, etc.). This is a hands-on technical leadership role managing a team of engineers building and scaling the event-driven services that form Klaviyo's messaging pipeline.
You will own critical components of the delivery infrastructure, drive technical strategy for a mission-critical system processing hundreds of millions of messages daily at scale, and help navigate multi-channel compliance, regulatory requirements, and global deliverability challenges. Key responsibilities include leading a team of engineers on highly-available, low-latency, high-throughput services; owning and architecting key pipeline components from ingestion through delivery and status tracking; designing scalable backend services in Python and/or Go built for real-time, unpredictable load; and making architectural tradeoffs on backpressure, delivery guarantees, and failure handling.
You will collaborate across product, deliverability, compliance, and platform teams to expand messaging capabilities, drive best practices in system design and code quality, mentor and grow engineers, and contribute to a high bar across the broader Messaging Infra organization. The role requires breaking down ambiguous technical problems into concrete, actionable plans and leveraging AI to improve team workflows and systems.
Required qualifications: 8+ years of engineering experience building and operating online, real-time distributed systems where latency and timely processing matter to customers or downstream services; proven ownership of production services measured in messages/events per second with concrete knowledge of p95/p99 latency under real traffic; hands-on experience with streaming or message-queue platforms (Kafka, Pulsar, RabbitMQ, SQS, or similar) as core system components; and demonstrated expertise architecting distributed, event-driven systems at scale with understanding of tradeoffs in keeping high-throughput pipelines reliable.