SlipstreamJobsFresh Startup & VC-Backed Jobs

Staff Site Reliability Engineer

IonQ - Santa Clara, CA, United States - Hybrid - posted 2026-09-17

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

Salary: USD 152,000 - 228,000 / annual

IonQ is the world's leading quantum computing platform and merchant supplier, delivering integrated quantum solutions across computing, networking, sensing, and security. The company achieved a world record in quantum computing performance in 2025 with 99.99% two-qubit gate fidelity and serves customers including Amazon Web Services and AstraZeneca. The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components. The Site Reliability Engineering discipline owns production reliability, service-level objectives, observability architecture, backup and disaster recovery, incident response, and resilience. As Staff Site Reliability Engineer, you will set the technical direction for reliability across regions and services. You own the reliability strategy, define standards and mechanisms that guide production operations, and raise the bar through design leadership, operational discipline, and mentorship. You remain deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, building resilience and disaster-recovery automation, improving reliability of stateful and streaming platforms, and creating AI Ops workflows for triage, remediation, and self-healing. Key responsibilities include: - Own service-level objectives, error budgets, and production reliability outcomes end to end - Design and operate the observability stack so production services are fully instrumented - Define and manage service-level objectives and drive corrective action when services consume error budgets unsafely - Design and execute chaos experiments and validate failure modes are covered by tested safeguards - Define incident processes and serve as incident commander for highest-severity incidents - Establish and manage on-call rotations and escalation paths with continuous coverage - Own disaster-recovery testing and failover validation against defined recovery objectives - Co-own cloud security posture management, runtime vulnerability detection, and configuration-compliance monitoring with DevSecOps - Own reliability of stateful and streaming services, capacity planning, and autonomous agents for triage, predictive alerting, remediation, and self-healing - Mentor engineers at different seniority levels and set standards adopted across teams The role is based in Santa Clara, CA with the option to work a few days a week remotely. Travel up to 25%. Requirements: - 7+ years of production engineering experience with recent hands-on reliability work - Hands-on, recent experience operating large-scale, fault-tolerant production systems on AWS or GCP - Observability ownership: has instrumented production systems and governed service-level objectives and error budgets, not only installed dashboards - Resilience practice: has designed and executed failure experiments or disaster-recovery exercises with real failover validation - Incident command: has personally commanded serious SEV1/SEV2 incidents and driven root cause through to a systemic fix - Demonstrated ownership of reliability outcomes with measurable results, such as availability, mean time to recovery, and error-budget adherence - Evidence of multi-team technical leadership through standards, review, coaching, and mechanisms adopted beyond one service or team Preferred Qualifications: - Proven production experience with cloud security posture management, runtime vulnerability detection, and workload protection across cloud and distributed environments - Strong experience prioritizing risk using identity, workload, and exposure-path context - Experience with autonomous remediation and self-healing workflows powered by AIOps, including Amazon Bedrock Agent Core or equivalent agentic automation frameworks - Hands-on experience in capacity management, resource rightsizing, efficiency engineering, and practical cost optimization based on FinOps principles - Experience with load-balancing design and operations, including health-based failover, global traffic management, and performance optimization - Experience with AI traffic management via an LLM gateway, including request routing, policy enforcement, rate limiting, model fallback, latency optimization, cost controls, and observability - Ability to connect networking, security, and reliability considerations into cohesive platform design decisions

Similar roles