SlipstreamJobsFresh Startup & VC-Backed Jobs

Senior Staff Service Reliability and Operational Intelligence Engineer

IonQ - Santa Clara, CA, United States - In-office - posted 2026-08-15

Apply on the company site

SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.

IonQ is the world's leading quantum computing platform and merchant supplier, delivering integrated quantum solutions across computing, networking, sensing, and security. The company achieved a world record in quantum computing performance in 2025 with 99.99% two-qubit gate fidelity and serves major customers including Amazon Web Services and AstraZeneca. The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites. The Service Reliability and Operational Intelligence discipline ensures the platform remains stable and resilient, with focus on service continuity and seamless customer experience. As a Senior Staff Service Reliability and Operational Intelligence Engineer, you will define the technical direction for reliability across regions and services. You own the reliability strategy, establish the standards and mechanisms that guide production operations, and elevate excellence through design leadership, operational discipline, and mentorship. You stay deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, and building resilience and disaster-recovery automation. Key responsibilities include owning the technical strategy and multi-year roadmap for operational excellence across development, pre-production, and production environments; defining and governing the New Service Introduction framework with mandatory architecture, security, resilience, capacity, observability, and release-readiness reviews; establishing organization-wide service ownership standards covering service catalogs, accountable owners, dependency maps, runbooks, support models, and on-call readiness; and leading the architecture and evolution of the shared observability platform with consistent standards for logs, metrics, distributed traces, and profiles. The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system. Up to 25% travel required.

Similar roles