SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Braze is seeking a Senior Site Reliability Engineer to own and operate the company's critical ingress infrastructure serving 3.3 billion monthly active users. In this role, you will architect and manage NGINX and Kubernetes-based ingress fleets that handle hundreds of billions of data points monthly and billions of messages daily.
You will lead the design and operation of high-performance NGINX routing and proxying layers, managing massive real-time API ingestion traffic. You'll own automated scaling routines for high-throughput services, leveraging RED metrics and Horizontal Pod Autoscalers to handle extreme traffic spikes seamlessly.
Working closely with product engineering teams, you'll translate feature requirements into resilient, highly available, and scalable technology stacks. You'll establish Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for API services, conduct systems design and capacity planning, and help teams balance rapid feature deployment with stability through error budget management.
You'll participate in PagerDuty on-call rotations, focusing not just on resolving alerts but on updating runbooks and preventing incident recurrence. You'll lead root-cause analysis and blameless retrospectives for availability and performance incidents, translating operational learnings into permanent system improvements.
Required: 5+ years as a DevOps or SRE in high-scale production environments. Deep hands-on experience with NGINX configuration and troubleshooting under heavy loads. In-depth Kubernetes proficiency including cluster networking, orchestration, and scheduling. Strong Linux/Unix internals knowledge (disk I/O, memory, TCP/IP, process management). Strong programming skills in Ruby, Go, Python, or Java. Experience with Infrastructure as Code (Terraform, Ansible, Chef). Systems thinking and understanding of distributed systems design, failure modes, and cascading effects.