SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, Infrastructure Engineer

Vapi - San Francisco, CA, USA - In-office - posted 2026-09-12

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 280,000 - 314,000 / annual

Vapi is a voice AI platform powering 1 billion calls for enterprise customers like Amazon Ring, Intuit, ServiceTitan, and New York Life. The company has raised $72M from top-tier investors including Y Combinator, Kleiner Perkins, Bessemer, and Peak XV. Vapi's real-time voice platform depends on core infrastructure spanning compute, storage, networking, and telephony. As usage scales, reliability must be architected into every system. You will join a five-person Infrastructure team as a dedicated Site Reliability Engineer, bringing a software-first mindset to build tooling, automation, and observability improvements that turn operational lessons into lasting engineering gains. In your first 30 days, you will learn Vapi's architecture, production environment, incident history, and reliability practices while building context with Infrastructure and product engineering teams. You'll contribute an initial operational or reliability improvement. By 60 days, you will own a reliability workstream—whether observability, incident response, capacity planning, performance optimization, or production automation. Your focus will be reducing manual work and improving how the team detects, understands, and responds to failures. By 90 days, you will become the go-to owner for a meaningful part of Vapi's reliability surface, delivering a durable improvement to failure prevention or recovery, and proposing a roadmap for future reliability investments. Vapi's culture emphasizes craft, follow-through, urgency, direct customer feedback, shared ownership, and kind directness. The founding team (Jordan and Nikhil, both Canadian) has built a company where 70% of employees are previous founders, creating a strong ownership mentality. REQUIREMENTS: - Senior or Staff-level software engineer with meaningful SRE, production engineering, or infrastructure experience in distributed systems - Production-quality software writing skills; experience building reliability tooling or automation (not just operational process) - Deep experience with observability, incident response, failure analysis, capacity planning, and production system health practices - Comfort with Kubernetes, networking, and cloud infrastructure; ability to debug across application and infrastructure boundaries - Clear reasoning about failure modes and ability to balance reliability investments with product and engineering velocity - Strong plus: experience with real-time networking or telephony, Envoy, Postgres, Redis, Kafka, Aurora, ClickHouse, or Google-style SRE environments

Similar roles