SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Site Reliability Engineer

Parasail - Remote - Remote - posted 2026-09-29

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Parasail is building an enterprise-grade inference cloud for open-weight models, delivering fast, reliable, and economical AI inference at scale. The company has raised $32M in Series A funding and is scaling beyond trillions of tokens per day. As a Senior Site Reliability Engineer, you will own infrastructure across Parasail's global GPU fleet and build the systems that ensure reliability across the entire stack. You'll work directly with infrastructure, platform, and inference engineers in a flat organization, taking responsibility for how systems perform in production. Key responsibilities include: - Scaling and improving Kubernetes infrastructure for provisioning, networking, storage, and service deployment across multiple providers and regions - Designing isolation, failover, and recovery mechanisms to minimize the impact of hardware and infrastructure failures on customers - Building software and automation to handle capacity expansion, deployments, and maintenance, reducing manual work and improving safety - Developing observability and diagnostics tools that reveal bottlenecks, surface failures, and enable rapid response - Owning the production feedback loop: responding to incidents, conducting root-cause analysis, and strengthening systems based on learnings - Collaborating across the stack to improve performance, utilization, security, and reliability as inference demand grows You'll work close to hardware, deep in distributed systems, and alongside engineers optimizing the inference stack. This is a small team tackling problems at substantial scale, with meaningful architecture decisions and direct production impact. REQUIREMENTS: - Production experience building and operating infrastructure or distributed systems with real ownership of reliability - Strong Linux fundamentals and practical knowledge of networking, storage, and containers - Hands-on experience running Kubernetes in production - Ability to write maintainable software and automation for infrastructure problems - Systematic debugging approach across application, cluster, network, and hardware boundaries - Good judgment about when to move quickly, simplify, or prioritize reliability - Initiative to drive problems from investigation through implementation with close team collaboration NICE TO HAVE: - Multi-region, multi-provider, or bare-metal infrastructure experience - Familiarity with GPUs, model serving, or inference systems (vLLM, SGLang) - Infrastructure as code, CI/CD, observability, or automated recovery experience - Experience building highly available services, multi-tenant platforms, or distributed data systems

Similar roles