SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer I

Braze - New York, NY, United States - In-office - posted 2026-08-26

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Braze is seeking a Senior Site Reliability Engineer to own and operate the company's ingress infrastructure and API ingestion layers—the critical entry point serving 3.3 billion monthly active users and billions of daily messages. This role focuses on the Ruby on Rails monolith and Go API services, requiring deep expertise in NGINX and Kubernetes to manage extraordinary scale. You will architect and operate high-performance NGINX routing and ingress controller layers handling massive real-time API traffic. You'll own and expand automated scaling routines using RED metrics, Horizontal Pod Autoscalers, and custom policies to handle extreme traffic spikes seamlessly. Working directly with product engineering teams, you'll translate feature requirements into resilient, highly available, and scalable technology stacks. Key responsibilities include establishing Service Level Indicators (SLIs) and Service Level Objectives (SLOs) for API services, conducting systems design and capacity planning to meet enterprise-grade SLAs, and participating in PagerDuty on-call rotations. You'll lead root-cause analysis and blameless retrospectives for incidents, translating operational learnings into permanent system improvements. Required: 5+ years as a DevOps or SRE in high-scale production environments. Deep hands-on experience with NGINX configuration and troubleshooting under heavy loads, in-depth Kubernetes administration and cluster networking, strong Linux/Unix internals knowledge (disk I/O, memory, TCP/IP, process management), and strong programming skills in Ruby, Go, Python, or Java. Infrastructure as Code experience (Terraform, Ansible, Chef) and systems thinking across distributed systems are essential.

Similar roles