← Back to results

Site Reliability Engineer, Provider Operations

Join OpenRouter as the first AI Inference SRE to ensure operational health of provider supply in a fully remote role.

Location
United States
Compensation
Not disclosed
Level
mid
Type
full time · Remote

Posted by employer 1 day ago

First seen on Joblaze 3 hours ago

Last verified on the company career page 3 hours ago

Apply at OpenRouter → Save job Scanned from openrouter.ai

What you'll build

  • Build and own monitoring for every provider and endpoint
  • Improve detection of degraded endpoints
  • Own on-call for provider incidents
  • Turn telemetry into scorecards and SLO reporting
  • Build continuous canaries and evals

Must have

  • 4+ years in SRE, production engineering, or infrastructure roles
  • Strong with observability tooling and practice
  • Capable software engineer who prefers writing tools
  • Experienced with distributed systems failure modes
  • Calm, clear incident commander

Nice to have

  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company
  • Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel
  • Background in routing, load balancing, or traffic management systems
  • Experience with evals or synthetic monitoring for ML systems

AI in the day-to-day

OpenRouter routes requests and manages large language models across providers.

Requirements

Experience
4+ years

Not disclosed in this posting: compensation, visa sponsorship.

Joblaze summary

In this role, the Site Reliability Engineer will focus on ensuring the operational health of AI provider endpoints, monitoring performance metrics, and implementing automated solutions to enhance reliability. Key skills include proficiency in observability tools, software engineering in TypeScript or Python, and a solid understanding of distributed systems. This position is ideal for someone with over four years of experience in SRE or production engineering, particularly in high-traffic environments. OpenRouter's emphasis on managing complex AI routing demands a candidate who can effectively handle incidents and improve system resilience.

Joblaze insights

  • Listed today — first seen on Joblaze October 11, 2026. Last confirmed on OpenRouter's careers page October 11, 2026.
  • Python appears in 58.3% of 120 comparable mid devops/sre roles in United States; TypeScript appears in 5.8% of 120 comparable mid devops/sre roles in United States.

Quick facts

Is the Site Reliability Engineer, Provider Operations role remote?
Yes — OpenRouter lists this as a fully remote position.
How much experience is required?
At least 4 years of relevant experience for this Site Reliability Engineer, Provider Operations role.
What's the tech stack?
Joblaze extracted these technologies from the posting: ClickHouse, Cloudflare Workers, GCP, PostgreSQL, Python, TypeScript.
What seniority level is this role?
OpenRouter targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Site Reliability Engineer, Provider Operations role at OpenRouter.

From the original posting

About OpenRouter

OpenRouter is the leading AI routing and infrastructure layer that enterprises use to access, manage, and optimize the best large language models across providers—without lock-in, capacity constraints, or unnecessary cost. We power the most advanced AI teams in the world by giving them the flexibility to move fast, scale confidently, and stay future-proof as models evolve.

About the Role

OpenRouter routes almost a billion requests and more than 20 trillion tokens a day, across 80+ providers and thousands of endpoints. Every one of those providers can degrade, rate-limit, change behavior, or go down without warning. Our customers count on us to absorb that chaos so their apps never notice.

We're hiring our first AI Inference SRE to own the operational health of our provider supply. You'll make sure every endpoint we route to is fast, correct, and available, and that we detect and route around problems before customers do. You'll sit on the Provider Operations team, reporting to the Provider Operations Manager.

What You'll Do

  • Provider health and observability. Build and own monitoring for every provider and endpoint: latency, throughput, error rates, uptime, and output correctness. Set SLOs per provider tier and alert on them.

  • Detection and failover. Improve how quickly we detect degraded endpoints, and work with the routing team so traffic shifts away from them automatically.

  • Incident response. Own on-call for provider incidents: triage, mitigate, communicate with providers, run postmortems, and drive follow-ups to closure.

  • Provider accountability. Turn telemetry into scorecards and SLO reporting that providers act on, and be the technical escalation point when a provider's endpoint is misbehaving.

  • Quality regression detection. Build continuous canaries and evals that catch silent regressions (quantization changes, broken tool calling, truncated streams, pricing or usage-reporting mismatches), not just outright downtime.

  • Automate the toil. Replace manual provider-ops work (disabling endpoints, capacity changes, deprecations, rate-limit tuning) with safe, auditable tooling.

  • Capacity and launch readiness. Build tooling to load-test endpoints before big launches so day-zero traffic doesn't take them down.

About You

  • 4+ years in SRE, production engineering, or infrastructure roles running high-traffic, customer-facing systems.

  • Strong with observability tooling and practice: metrics, tracing, logs, SLOs/error budgets, alerting that is always actionable.

  • Capable software engineer who prefers writing tools over executing runbooks. TypeScript and/or Python.

  • Experienced with distributed systems failure modes: timeouts, retries, backpressure, partial outages, noisy neighbors.

  • Calm, clear incident commander who communicates well with external partners under pressure.

  • Understands, or is eager to learn deeply, how LLM inference is served: streaming, tool calling, prompt caching, throughput/latency tradeoffs, and how provider APIs differ.

Nice to Have

  • Experience at an inference provider, model lab, GPU cloud, or API gateway/CDN company.

  • Experience with our stack: TypeScript, Cloudflare Workers, Postgres, ClickHouse, GCP, Vercel.

  • Background in routing, load balancing, or traffic management systems.

  • Experience with evals or synthetic monitoring for ML systems.

Standard company text repeated across OpenRouter's postings is omitted here.

Similar positions

OpenRouter
OpenRouter
Software Engineer, Trust & Safety
OpenRouter · Remote (US)
OpenRouter
Applied AI Engineer
OpenRouter · Remote (US)
OpenRouter
Software Engineer, Platform
OpenRouter · Remote (US)