SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Braze is seeking a Senior Site Reliability Engineer to own and operate the company's critical ingress infrastructure serving 3.3 billion monthly active users. In this role, you will architect and operate high-performance NGINX routing and Kubernetes ingress controller layers that manage massive real-time API ingestion traffic—the main entry point for how the internet communicates with Braze's platform.
You will lead the design and operation of ingress fleets, configure and tune automated scaling routines for high-throughput API services using RED metrics and Horizontal Pod Autoscalers, and partner directly with product engineering teams to translate feature requirements into resilient, highly available technology stacks. You'll establish Service Level Indicators and Objectives for API services, conduct systems design and capacity planning to meet enterprise-grade SLAs, and participate in on-call rotations to resolve incidents and prevent recurrence through blameless retrospectives and root-cause analysis.
This is a hands-on technical role requiring deep expertise in NGINX and Kubernetes administration. You will work at extraordinary scale, handling hundreds of billions of data points monthly and billions of messages daily. The position is embedded with product teams, giving you direct influence over how Braze's foundational infrastructure evolves.
Required: 5+ years as a DevOps or SRE in high-scale production environments; deep hands-on NGINX experience; in-depth Kubernetes proficiency including cluster networking and orchestration; strong Linux/Unix internals knowledge; programming skills in Ruby, Go, Python, or Java; Infrastructure as Code experience (Terraform, Ansible, Chef); and systems thinking across distributed systems. You should excel at translating operational learnings into permanent improvements and thrive in a collaborative, fast-paced environment.