SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Braze is seeking a Senior Site Reliability Engineer to own and operate the company's NGINX and Kubernetes ingress infrastructure that handles massive, real-time API ingestion traffic serving over 3.3 billion monthly active users. In this role, you will architect and operate high-performance ingress fleets, configure and tune NGINX routing and proxying layers, and own automated scaling routines for high-throughput API services using RED metrics, Horizontal Pod Autoscalers, and customized scaling policies.
You will partner directly with product engineering teams to translate feature requirements into resilient, highly available, and scalable technology stacks. This includes establishing meaningful Service Level Indicators (SLIs) and Service Level Objectives (SLOs), conducting in-depth systems design and capacity planning, and managing error budgets to balance rapid feature deployment with stability. You will participate in PagerDuty on-call rotations, lead root-cause analysis and blameless retrospectives for availability and performance incidents, and translate operational learnings into permanent system improvements.
Required qualifications include 5+ years of professional experience as a DevOps or Site Reliability Engineer in high-scale production environments. You must have deep hands-on experience with NGINX configuration and troubleshooting under heavy traffic loads, in-depth proficiency with Kubernetes administration and cluster networking, and excellent OS-level understanding of Linux/Unix internals including disk I/O, memory allocation, TCP/IP networking, and process management. Strong programming and scripting skills are essential—Ruby and/or Go preferred, or equivalent languages like Python or Java—to build custom automated tools and platform frameworks. Experience with Infrastructure as Code technologies such as Terraform, Ansible, or Chef is required, along with strong conceptual understanding of systems design, failure modes, and cascading effects across distributed systems.