← Back to results

Staff Service Reliability and Operational Intelligence Engineer

Define the technical direction for reliability and operational intelligence at IonQ, focusing on service continuity and customer experience.

Location
Santa Clara, California, United States
Compensation
Not disclosed
Level
staff
Type
full time · Hybrid

Posted by employer 1 week ago

First seen on Joblaze 6 days ago

Last verified on the company career page 3 hours ago

Apply at IonQ → Save job Scanned from ionq.com

Skills & Technologies

What you'll build

  • Shape the technical strategy for operational excellence
  • Define and govern the New Service Introduction framework
  • Lead the architecture of the shared observability platform
  • Own the reliability governance model for production services
  • Advance incident-management maturity

Must have

  • 8+ years of production engineering experience
  • Recent experience designing large-scale production systems
  • Deep understanding of distributed systems and cloud infrastructure
  • Demonstrated ownership of observability architecture
  • Experience commanding SEV1 or SEV2 incidents

Nice to have

  • Experience prioritizing operational risk
  • Experience designing AI Ops capabilities
  • Hands-on experience with autonomous remediation
  • Experience with capacity optimization
  • Experience with progressive-delivery techniques

Practical constraints

  • Travel: Up to 25%

Requirements

Experience
8+ years

Not disclosed in this posting: compensation, visa sponsorship.

Benefits

401k Match Unlimited PTO Equity/Stock Options Health Insurance Parental Leave

Joblaze summary

In this role, the Staff Service Reliability and Operational Intelligence Engineer at IonQ is responsible for shaping the technical strategy for reliability across various services, ensuring system stability and resilience. Key skills include expertise in cloud infrastructure, distributed systems, and observability architecture, with a strong emphasis on automation and incident management. This position is suited for seasoned professionals with extensive experience in production engineering and a proven track record in operational excellence. The role involves hands-on leadership during critical incidents and the development of AI Ops capabilities to enhance service reliability.

Joblaze insights

  • Listed 6 days ago — first seen on Joblaze September 19, 2026. Last confirmed on IonQ's careers page September 25, 2026.
  • Kubernetes appears in 64.9% of 77 comparable staff devops/sre roles in United States; GCP appears in 31.2% of 77 comparable staff devops/sre roles in United States.

Quick facts

Is the Staff Service Reliability and Operational Intelligence Engineer role remote?
It's hybrid — IonQ expects some on-site time in Santa Clara, California, United States.
How much experience is required?
At least 8 years of relevant experience for this Staff Service Reliability and Operational Intelligence Engineer role.
Where is the role based?
IonQ is hiring for this position in Santa Clara, California, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, GCP, Go, Kubernetes, Python.
What seniority level is this role?
IonQ targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Staff Service Reliability and Operational Intelligence Engineer role at IonQ.

From the original posting

About IonQ:

Location: This role is based at our Santa Clara, CA office, with the option to work a few days a week remotely.
Travel: Up to 25%
Job ID:
1875

The Role:

The Platform Engineering team builds, secures, and operates scalable infrastructure for cloud-managed SaaS products with on-premises components deployed at customer sites.

The Service Reliability and Operational Intelligence discipline ensures the platform remains stable and resilient, with focus on service continuity and seamless customer experience. It owns production reliability and resilience, observability architecture, service-level objectives, incident response, and implementation of AIOps workflows for triage, remediation, and self-healing.

As a Staff Service Reliability and Operational Intelligence Engineer, you define the technical direction for reliability across regions and services. You own the reliability strategy, establish the standards and mechanisms that guide production operations, and elevate excellence through design leadership, operational discipline, and mentorship. You stay deeply hands-on by designing and operating observability platforms, defining and governing SLO programs, leading high-severity incident response, and building resilience and disaster-recovery automation.

The work is driven by observability and automation, with a focus on detecting and fixing issues before customers are affected and using every incident to improve the system.

Responsibilities:

  • Shape the technical strategy and multi-year roadmap for operational excellence and production readiness across development, pre-production, and production environments.
  • Define and govern the New Service Introduction framework, including mandatory architecture, security, resilience, capacity, observability, supportability, and release-readiness reviews before services enter production.
  • Establish organization-wide service ownership standards covering service catalog records, accountable owners, dependency maps, runbooks, support models, escalation paths, recovery objectives, and on-call readiness.
  • Lead the architecture and evolution of the shared observability platform, establishing consistent standards for logs, metrics, distributed traces, and profiles across production systems.
  • Define standards for dashboards, alert policies, synthetic monitoring, telemetry quality, retention, sampling, cardinality, and cost controls.
  • Own the reliability governance model for production services, including SLIs, SLOs, error budgets, and escalation mechanisms.
  • Connect service-health signals to customer and business impact, enabling early anomaly detection, service-degradation prevention, and rapid isolation of end-user-impacting events.
  • Advance incident-management maturity through consistent severity classification, incident command, stakeholder and executive communications, automated evidence collection, and coordinated response to high-severity incidents.
  • Establish blameless post-incident review practices, ensure remediation actions are tracked to completion, and drive systemic fixes for recurring failure modes.
  • Lead operational capacity and efficiency management, including demand forecasting, cloud and Kubernetes capacity, performance testing, scaling thresholds, headroom policies, resource rightsizing, and capacity-risk reviews.
  • Design and govern AI Ops capabilities for event correlation, alert-noise reduction, predictive detection, probable root-cause analysis, autonomous triage, assisted remediation, and controlled self-healing.
  • Deliver secure AI-agent workflows across observability platforms, service catalog, Jira, Confluence, source control, and CI/CD.
  • Improve on-call effectiveness through sustainable rotation design, operational-readiness standards, escalation policies, diagnostic automation, alert-quality management, and reliable follow-the-sun handoffs.
  • Provide hands-on technical leadership during major incidents, complex reliability investigations, architectural reviews, resilience exercises, and critical service launches.
  • Use operational data, incident trends, service-level performance, capacity signals, change outcomes, and automation effectiveness to prioritize continuous improvement.

Requirements:

  • 8+ years of production engineering, site reliability engineering, platform engineering, or cloud operations experience, including recent hands-on reliability work.
  • Recent experience designing and operating large-scale, fault-tolerant production systems on AWS or GCP.
  • Deep understanding of distributed systems, cloud infrastructure, Kubernetes, networking, CI/CD, and production failure modes.
  • Demonstrated ownership of observability architecture, including instrumentation of production systems and governance of metrics, logs, traces, SLIs, SLOs, and error budgets.
  • Proven experience establishing reliability and operational-readiness standards for business-critical services.
  • Hands-on experience designing and executing failure experiments, disaster-recovery exercises, and validated service failovers.
  • Experience personally commanding SEV1 or SEV2 incidents, coordinating technical and executive communications, and driving root causes through to systemic remediation.
  • Demonstrated ownership of measurable reliability outcomes such as availability, latency, MTTR, change-failure rate, alert quality, and error-budget adherence.
  • Experience with capacity forecasting, performance testing, scaling strategies, and cloud and Kubernetes resource management.
  • Strong software engineering and automation skills using languages such as Python or Go, infrastructure as code, and modern delivery toolchains.
  • Evidence of multi-team technical leadership through architecture reviews, standards, coaching, and mechanisms adopted beyond a single service or team.
  • Ability to influence cross-functional stakeholders and deliver complex initiatives without relying on direct management authority.

Preferred Qualifications:

  • Experience prioritizing operational risk using identity, workload, dependency, and exposure-path context to focus remediation on issues with material customer or business impact.
  • Experience designing AI Ops capabilities for anomaly detection, event correlation, predictive alerting, root-cause analysis, and operational noise reduction.
  • Hands-on experience with autonomous remediation and self-healing workflows using Amazon Bedrock AgentCore or comparable agentic automation frameworks.
  • Experience integrating governed AI agents with operational platforms such as Jira, Confluence, source control, CI/CD, service catalog, and observability systems.
  • Practical experience with capacity optimization, resource rightsizing, efficiency engineering, telemetry cost management, and FinOps principles.
  • Experience designing and operating load-balancing solutions, health-based failover, global traffic management, and performance optimization for highly available services.
  • Ability to integrate networking, security, resilience, performance, and operability requirements into cohesive platform architecture decisions.
  • Experience with progressive-delivery techniques such as canary deployments, blue-green deployments, automated rollback, and feature-flag governance.
  • Experience establishing sustainable global on-call models and follow-the-sun operational practices.


The total compensation package includes base, bonus, equity, and a range of benefit options found on our career site.

If this role has a commission structure, the compensation range below just reflects the base compensation range.

Wage Transparency:
$152,000—$228,000 USD

Compensation will vary based on individual factors such as education, qualifications, and experience of the final candidate(s), specific office location, and calibration against relevant market data and internal team equity. Posted base salary figures are subject to change as new market data becomes available. Our benefits include comprehensive medical, dental, and vision plans, matching 401(k), unlimited PTO and paid holidays, parental/adoption leave, legal insurance, and a home technology stipend. Details of participation in these benefit plans will be provided when a candidate receives an offer of employment.

Standard company text repeated across IonQ's postings is omitted here.

Similar positions

IonQ
Staff Site Reliability Engineer
IonQ · Santa Clara, California, United States
IonQ
Staff DevOps Engineer
IonQ · Santa Clara, California, United States
IonQ
Staff Distributed Systems Engineer
IonQ · Santa Clara, California, United States
IonQ
Senior Staff Distributed Systems Engineer
IonQ · Santa Clara, California, United States
IonQ
Senior Distributed Systems Engineer
IonQ · Santa Clara, California, United States