SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Parasail - San Mateo, CA, USA - In-office - posted 2026-09-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Parasail is building an enterprise-grade inference cloud for open-weight models, delivering fast, reliable, and economical AI inference at scale. The company pools GPU capacity from providers worldwide and optimizes workload placement through a single OpenAI-compatible API. With $32M in Series A funding, Parasail is scaling beyond trillions of tokens per day. As a Senior Site Reliability Engineer, you will own infrastructure across Parasail's global GPU fleet and build systems that ensure reliability across the entire stack. You'll work in a flat organization alongside infrastructure, platform, and inference engineers, taking direct ownership of how systems perform in production. Key responsibilities include: - Scaling and improving Kubernetes infrastructure for provisioning, networking, storage, and service deployment across multiple providers and regions - Designing isolation, failover, and recovery mechanisms to minimize the impact of hardware and infrastructure failures on customers - Building software and automation to handle capacity expansion, deployments, and maintenance, reducing manual work and improving safety - Developing observability and diagnostics tools that reveal bottlenecks, surface failures, and enable rapid response - Owning the production feedback loop by responding to incidents, conducting root-cause analysis, and strengthening systems based on learnings - Collaborating across the stack to improve performance, utilization, security, and reliability as inference demand grows You'll work close to hardware, deep in distributed systems, and help build the foundation for the next stage of AI infrastructure. The systems you build will directly determine how reliably and efficiently customers can run AI in production. REQUIREMENTS: - Production experience building and operating infrastructure or distributed systems with real ownership of reliability - Strong Linux fundamentals and practical knowledge of networking, storage, and containers - Hands-on experience running Kubernetes in production - Ability to write maintainable software and automation for infrastructure problems - Systematic debugging approach across application, cluster, network, and hardware boundaries - Good judgment about when to move quickly, when to simplify, and where reliability matters most - Initiative to take problems from investigation through implementation with close team collaboration NICE TO HAVE: - Experience with multi-region, multi-provider, or bare-metal infrastructure - Familiarity with GPUs, model serving, or inference systems (vLLM, SGLang) - Infrastructure as code, CI/CD, observability, or automated recovery experience - Experience building highly available services, multi-tenant platforms, or distributed data systems

Similar roles