SlipstreamJobsFresh Startup & VC-Backed Jobs

Member of Technical Staff, Site Reliability Engineer

Inferact - San Francisco, CA, USA - Hybrid - posted 2026-08-20

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 200,000 - 400,000 / annual

Inferact, founded by the creators and core maintainers of vLLM, is building the world's AI inference engine. The company sits at the intersection of models and hardware, focused on making inference cheaper and faster at production scale. You'll join as a Site Reliability Engineer to ensure vLLM-powered inference systems are reliable, observable, and operationally simple. This role bridges engineering and infrastructure, with direct impact on the reliability and availability of AI inference systems serving production traffic. Key responsibilities include defining SLOs and SLIs, improving monitoring and alerting, strengthening incident response processes, driving post-mortems, and reducing operational risk before it reaches users. You'll think about failure modes before launch, design systems that are easier to operate, and turn incidents into durable improvements. Required experience: production systems operation with meaningful traffic or infrastructure criticality; deep knowledge of SLOs, SLIs, error budgets, alerting, and incident response; hands-on experience fighting major production incidents including mitigation and root cause analysis; strong Linux, networking, systems debugging, observability, and distributed systems fundamentals; ability to design operationally simple systems; strong programming or scripting in Python, Go, Bash, or similar. Preferred: ML infrastructure or inference systems experience; GPU workloads or Kubernetes platforms; high-scale backend services; building observability systems with metrics, logs, traces, dashboards, and runbooks; Kubernetes, Docker, Terraform, cloud infrastructure, service meshes, or CI/CD experience; incident review culture and prevention-oriented engineering; partnership with engineering teams on service design and operational readiness. Bonus: ownership of high-throughput, latency-sensitive, or mission-critical production systems; AI inference, model serving, GPU clusters, or distributed serving infrastructure support; automation that reduced toil or prevented repeat incidents; incident response leadership for severe outages; practical SLOs, dashboards, alerts, or release gates that improved reliability. Based in San Francisco with remote consideration for exceptional US-based candidates.

Similar roles