SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
Fireworks is a Series D AI infrastructure platform enabling companies to build, train, and serve specialized AI models on their own data. Founded by the PyTorch team and backed by AMD, NVIDIA, Sequoia, and others, Fireworks powers production inference and training for open models, serving over 40 trillion tokens daily.
As a Member of Technical Staff in Reliability Engineering, you will own the reliability posture of Fireworks' platform as it scales. This is a high-leverage individual contributor role focused on systems reliability across GPU scheduling, kernel performance, networking, storage, and Linux infrastructure.
Key responsibilities include:
- Define and drive adoption of SLOs, error budgets, and production readiness standards across engineering teams
- Own the reliability toolchain end-to-end: logging, telemetry pipelines, alerting, failure injection, load testing, and self-healing automation
- Identify and fix cross-system failure modes (retry amplification, timeout composition, unmapped dependencies) before customers encounter them
- Coordinate incident management, run blameless postmortems, and track follow-ups
- Reduce operational toil through automation to keep on-call sustainable as the platform grows
- Partner with cloud infrastructure, inference/training, performance, and product teams to ensure customer experience reliability
You will have broad visibility into the platform and latitude to focus your time where it drives the most impact. The role emphasizes influence without authority—getting teams to adopt standards through credibility and useful tooling rather than mandate.
Required: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC); 5+ years writing production-grade tools in Python, Go, C++, or Rust; hands-on experience operating Kubernetes, Terraform, and Docker in high-throughput production; familiarity with distributed systems, fault-tolerant design, and SLO/SLA management; ability to influence across teams; comfort with breadth over depth; Bachelor's or Master's in Computer Science or equivalent.
Preferred: Observability tools (Prometheus, Grafana, OpenTelemetry); GPU and ML infrastructure exposure; AI-assisted operations; open source contributions; startup agility.