SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Braze is seeking a Senior Site Reliability Engineer to own and operate the company's NGINX and Kubernetes ingress infrastructure, which serves as the critical entry point for billions of API requests daily from across the internet. You will architect, configure, and tune high-performance routing and proxying layers that handle massive real-time API ingestion traffic at extraordinary scale—serving 3.3 billion monthly active users, collecting hundreds of billions of data points monthly, and delivering billions of messages daily.
In this role, you will lead the design and operation of ingress fleets, own and expand automated scaling routines for high-throughput API services using RED metrics and Horizontal Pod Autoscalers, and partner directly with product engineering teams to translate feature requirements into resilient, highly available, and scalable technology stacks. You will establish meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for API services, conduct in-depth systems design and capacity planning, and help teams navigate error budgets to balance rapid feature deployment with stability.
You will participate in a PagerDuty on-call rotation, utilizing shifts not just to resolve structural alerts but to update runbooks and proactively prevent incidents. You will lead root-cause analysis and blameless retrospectives for availability and performance incidents, translating operational learnings into permanent system improvements.
Required: 5+ years as a DevOps or Site Reliability Engineer in high-scale production environments; deep hands-on experience with NGINX configuration, troubleshooting, and operation under heavy traffic; in-depth Kubernetes administration and cluster networking proficiency; excellent Linux/Unix internals knowledge (disk I/O, memory, TCP/IP, process management); strong programming/scripting skills in Ruby, Go, Python, or Java; experience with Infrastructure as Code (Terraform, Ansible, Chef); and strong systems thinking around distributed systems design, failure modes, and cascading effects.