SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Inference Reliability Engineer

Parasail - San Mateo, CA, United States - In-office - posted 2026-09-25

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Parasail is redefining AI infrastructure by enabling seamless deployment across a distributed network of GPUs, optimizing for cost, performance, and flexibility. The Senior Inference Reliability Engineer will own the end-to-end reliability and production performance of customer inference workloads, sitting at the intersection of inference platform engineering, LLM performance, and infrastructure reliability. You will ensure that customer endpoints meet expectations for availability, latency, throughput, quality, and cost. When an endpoint degrades, you will follow the problem across the entire serving path—from APIs, routing, scheduling, and autoscaling through model servers, GPUs, networking, and underlying infrastructure—and drive it through resolution. Key responsibilities include: • Owning production health of customer inference workloads, including availability, request success, time to first token, inter-token latency, throughput, and operational efficiency • Establishing clear service-level indicators, objectives, performance baselines, and escalation paths for production endpoints • Building telemetry, dashboards, alerts, and automated diagnostics to detect endpoint degradation before customers report it • Creating visibility across the full inference-serving path, including request queues, routing, scheduling, model servers, GPU utilization, networking, storage, and provider infrastructure • Leading investigation of complex latency, throughput, capacity, and reliability regressions across multiple system layers • Helping lead customer-impacting incidents and establishing effective operational practices for acknowledgement, diagnosis, recovery, and communication • Converting significant incidents into automated tests, safeguards, runbooks, capacity controls, anomaly detection, and platform improvements • Partnering with the LLM Performance team to validate that engine-level optimizations deliver measurable improvements in production • Analyzing workload behavior, capacity requirements, utilization, tail latency, and cost efficiency across heterogeneous GPU providers and hardware • Identifying recurring patterns across incidents, workloads, and customer escalations and translating findings into platform improvements This is a production systems role for an engineer who enjoys investigating ambiguous performance problems, building diagnostic tooling, and turning recurring incidents into durable platform improvements. Prior LLM-inference experience is valuable but not required; the role seeks someone with deep production systems experience who can quickly learn inference-specific technologies and metrics. Within the first six months, success means material endpoint regressions are increasingly detected before customers report them, engineers can quickly determine which system layer is responsible for production issues, customer-impacting incidents are acknowledged and resolved faster, production endpoints consistently meet defined objectives, and recurring failure modes are converted into automated detection and safeguards. REQUIREMENTS: • 5+ years of experience in production engineering, site reliability engineering, infrastructure engineering, distributed systems, ML infrastructure, database reliability, or performance engineering • Demonstrated ownership of a critical production service or workload • Experience diagnosing complex latency, throughput, capacity, or reliability problems across multiple system layers • Strong software-engineering ability beyond infrastructure configuration and CI/CD automation • Hands-on experience building observability, automation, diagnostic tooling, or production safeguards • Strong communication and technical leadership skills, including the ability to coordinate incident resolution across engineering teams • Experience with Kubernetes, Linux, networking, and cloud-native infrastructure • Demonstrated ability to learn unfamiliar systems and develop deep technical expertise • Experience with GPUs, ML infrastructure, model serving, vLLM, SGLang, Triton, TensorRT-LLM, or similar technologies is valuable but not required • Proficiency in languages such as Python, Go, Java, C++, or Rust

Similar roles