SlipstreamJobs tracks this role from the company's public career site. Apply directly on the employer's site.
IonQ is the world's leading quantum computing platform and merchant supplier, delivering integrated quantum solutions across computing, networking, sensing, and security. The company achieved 99.99% two-qubit gate fidelity in 2025, setting a world record in quantum computing performance. Customers and partners including Amazon Web Services and AstraZeneca use IonQ systems to accelerate innovation in drug discovery, materials science, financial modeling, logistics, cybersecurity, and defense.
The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites. The Service Reliability and Operational Intelligence discipline ensures the platform remains stable and resilient, with focus on service continuity and seamless customer experience.
As a Staff Service Reliability and Operational Intelligence Engineer, you will define the technical direction for reliability across regions and services. You will own the reliability strategy, establish standards and mechanisms that guide production operations, and elevate excellence through design leadership, operational discipline, and mentorship. The role stays deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, and building resilience and disaster-recovery automation.
Key responsibilities include: shaping technical strategy and multi-year roadmap for operational excellence across environments; defining and governing the New Service Introduction framework with mandatory architecture, security, resilience, and observability reviews; establishing organization-wide service ownership standards covering service catalog, dependency maps, runbooks, and on-call readiness; leading architecture and evolution of the shared observability platform with consistent standards for logs, metrics, traces, and profiles; defining standards for dashboards, alert policies, synthetic monitoring, and cost controls; owning reliability governance including SLIs, SLOs, and error budgets; connecting service-health signals to customer and business impact for early anomaly detection; advancing incident-management maturity through severity classification, incident command, and coordinated response to high-severity incidents; establishing blameless post-incident review practices and driving systemic fixes; leading operational capacity and efficiency management including demand forecasting and resource rightsizing; designing and governing AI Ops capabilities for event correlation, alert-noise reduction, predictive detection, and autonomous triage; delivering secure AI-agent workflows across observability platforms, service catalog, Jira, Confluence, source control, and CI/CD; improving on-call effectiveness through sustainable rotation design and alert-quality management; providing hands-on technical leadership during major incidents and architectural reviews; and using operational data to prioritize continuous improvement.
The role is based at the Santa Clara, CA office with the option to work a few days a week remotely. Travel up to 25% is required.
REQUIREMENTS:
- 8+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work
- Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP
- Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes
- Demonstrated ownership of observability architecture, including instrumentation of production systems and governance of metrics, logs, traces, SLIs, SLOs, and error budgets
- Proven experience establishing reliability and operational-readiness standards for business-critical services
- Hands-on experience designing and executing failure experiments, disaster-recovery exercises, and validated service failovers
- Experience personally commanding SEV1 or SEV2 incidents, coordinating technical and executive communications, and driving root causes through to systemic remediation
- Demonstrated ownership of measurable reliability outcomes such as availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence
- Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management
- Strong software engineering and automation skills using languages such as Python or Go, infrastructure as code, and modern delivery toolchains
- Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and mechanisms adopted beyond a single service or team
- Ability to influence cross-functional stakeholders and deliver complex initiatives without relying on direct management authority
PREFERRED QUALIFICATIONS:
- Experience prioritizing operational risk using identity, workload, dependency, and exposure-path context
- Experience designing AI Ops capabilities for anomaly detection, event correlation, predictive alerting, and root-cause analysis
- Hands-on experience with autonomous remediation and self-healing workflows using Amazon Bedrock AgentCore or comparable agentic automation frameworks
- Experience integrating governed AI agents with operational platforms such as Jira, Confluence, source control, CI/CD, service catalog, and observability systems
- Practical experience with capacity optimization, resource rightsizing, efficiency engineering, telemetry cost management, and FinOps principles
- Experience designing and operating load-balancing solutions, health-based failover, global traffic management, and performance optimization
- Ability to integrate networking, security, resilience, performance, and operability requirements into cohesive platform architecture decisions
- Experience with progressive-delivery techniques such as canary deployments, blue-green deployments, automated rollback, and feature-flag governance
- Experience establishing sustainable global on-call models and follow-the-sun operational practices