← Back to results

Site Reliability Engineer

Join Cognition as a Site Reliability Engineer to ensure production reliability and enhance platform engineering for AI-driven products.

Location
San Francisco
Compensation
$260k–$300k/yr
Level
senior
Type
full time

Posted by employer 1 month ago

First seen on Joblaze 2 days ago

Last verified on the company career page 11 hours ago

Apply at Cognition → Save job Scanned from cognition.ai

Skills & Technologies

What you'll build

  • Define and own SLOs, SLIs, and error budgets
  • Lead incident response and run blameless postmortems
  • Own deployment pipelines and release infrastructure
  • Manage cloud infrastructure through code
  • Model growth and forecast resource needs

Must have

  • Deep experience running production systems at scale
  • Strong software engineering fundamentals
  • Proficiency with cloud infrastructure
  • Experience building and owning CI/CD pipelines
  • Strong observability instincts
  • Comfort owning incidents end to end

Nice to have

  • Experience with developer-facing products or platforms

AI in the day-to-day

Building Devin, the first AI software engineer, and ensuring reliability for AI-driven products.

Not disclosed in this posting: years of experience, work arrangement, visa sponsorship.

Benefits

401k Match Equity/Stock Options Health Insurance

Joblaze summary

The Site Reliability Engineer at Cognition is responsible for ensuring the production reliability of user-facing products like Devin and Windsurf, managing incident response, and optimizing deployment infrastructure. This role requires strong skills in cloud infrastructure, CI/CD pipelines, and observability, alongside a solid foundation in software engineering. It is well-suited for experienced professionals who thrive in high-stakes environments and have a proactive approach to reliability. Cognition's small, elite team emphasizes ownership and a culture that integrates reliability into the development process.

Joblaze insights

  • Listed 2 days ago — first seen on Joblaze September 21, 2026. Last confirmed on Cognition's careers page September 23, 2026.
  • Salary band is above the typical range for DevOps/SRE roles (median ~$165,000).
  • Starts above 99% of 79 comparable senior devops/sre roles in United States that list Kubernetes we track (median $165,000 across 37 companies). See Kubernetes salary trends
  • Kubernetes appears in 69.2% of 185 comparable senior devops/sre roles in United States; Azure appears in 22.2% of 185 comparable senior devops/sre roles in United States.

Quick facts

What's the salary range?
Cognition lists $260,000–$300,000 for this role.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, Azure, GCP, Kubernetes, Terraform.
What seniority level is this role?
Cognition targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Site Reliability Engineer role at Cognition.

From the original posting

Role Mission

Devin and Windsurf are used by hundreds of thousands of developers every day. When something goes wrong, it goes wrong for all of them at once. This role exists to make sure that doesn't happen, and when it does, to make sure it's resolved faster than anyone expects.

You will own both the production reliability of our user-facing products and the platform engineering that lets our team ship quickly and confidently. That means SLOs, incident response, and on-call on one side, and CI/CD pipelines, deployment infrastructure, and developer tooling on the other. At Cognition, these are not separate jobs. The best SREs here understand that reliability is engineered in, not bolted on.

What You'll Accomplish

  • Production Reliability: Define and own SLOs, SLIs, and error budgets for Devin and Windsurf. Build the monitoring, alerting, and observability systems that give the team a clear, honest picture of service health at all times.

  • Incident Response and On-Call: Lead incident response with speed and clarity. Run blameless postmortems that turn outages into durable improvements. Build the runbooks and tooling that make on-call sustainable and effective.

  • Platform Engineering and CI/CD: Own the deployment pipelines, release infrastructure, and internal developer tooling that let the team ship fast without breaking things. Reduce toil systematically so engineers spend time on work that matters.

  • Infrastructure as Code: Manage cloud infrastructure through code. Build reproducible, auditable, version-controlled environments that scale with the product and eliminate configuration drift.

  • Capacity Planning and Performance: Model growth, forecast resource needs, and ensure the infrastructure stays ahead of demand. Profile and improve system performance before users feel it.

  • Security and Reliability as One: Treat security not as a separate concern but as a reliability requirement. Ensure that misconfigurations, vulnerabilities, and access failures are caught and remediated with the same urgency as outages.

  • Reliability Culture: Partner closely with product and engineering teams to build reliability in from the start. Be the person who catches the single point of failure in the architecture review before it becomes a page at 2am.

Exceptional Candidates Have Demonstrated

  • Deep experience running production systems at scale: SLOs, error budgets, on-call rotations, and incident command

  • Strong software engineering fundamentals; SRE at Cognition means writing real code, not just configuring tools

  • Proficiency with cloud infrastructure (AWS, GCP, or Azure), container orchestration (Kubernetes), and infrastructure as code (Terraform or equivalent)

  • Experience building and owning CI/CD pipelines and deployment infrastructure for fast-moving product teams

  • Strong observability instincts: knows how to instrument systems, build useful dashboards, and design alerts that surface signal without generating noise

  • A track record of reducing toil systematically through automation, not just working around it

  • Comfort owning incidents end to end: detection, triage, mitigation, resolution, and postmortem

  • Enough product empathy to understand what reliability means from a user's perspective, not just an infrastructure one

  • Experience with developer-facing products or platforms is a strong plus

Resources & Environment

  • Small, highly selective team shipping products used by hundreds of thousands of developers daily

  • High ownership and high trust: you'll set the reliability bar, not inherit someone else's standards

  • The environment rewards engineers who are proactive, systematic, and treat reliability as a craft, not a checklist

Compensation & Benefits

  • Base Salary: $260,000 - $300,000 + significant early-stage equity

  • Medical, Dental, Vision: Fully paid for you and your dependents

  • 401(k): Company match included

  • Perks: Private chef, cozy slippers, endless snacks, and more

Standard company text repeated across Cognition's postings is omitted here.

Similar positions

Cognition
Software Engineer, Infrastructure
Cognition · San Francisco
Cognition
Software Engineer
Cognition · San Francisco
Cognition
AI Support Engineer
Cognition · San Francisco
Cognition
Security Engineer
Cognition · San Francisco
Cognition
AI Support Engineer
Cognition · Tokyo