← Back to results

Site Reliability Engineer

Own the operational health of Gamma's backend platform while building automation and tooling to improve reliability for millions of users.

Location
San Francisco
Compensation
$230k–$310k/yr
Level
senior
Type
full time · Hybrid

Posted by employer 10 months ago

First seen on Joblaze 1 week ago

Last verified on the company career page 1 day ago

Apply at Gamma → Save job Scanned from gamma.app

Requirements

Experience
5+ years

Not disclosed in this posting: visa sponsorship.

Benefits

Equity/Stock Options Remote Work

Joblaze summary

In this role, the Site Reliability Engineer at Gamma is responsible for ensuring the operational health of the backend platform, focusing on building automation and observability tools that enhance system reliability. Key skills include deep AWS expertise, programming in Python or Go, and experience with infrastructure-as-code and observability solutions. This position is ideal for someone with over five years in site reliability or DevOps, particularly those who have scaled SaaS products to millions of users. Gamma fosters a collaborative in-office culture, emphasizing teamwork while allowing flexibility for focused work.

Joblaze insights

Quick facts

Is the Site Reliability Engineer role remote?
It's hybrid — Gamma expects some on-site time in San Francisco.
What's the salary range?
Gamma lists $230,000–$310,000 for this role.
How much experience is required?
At least 5 years of relevant experience for this Site Reliability Engineer role.
Where is the role based?
Gamma is hiring for this position in San Francisco.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, CloudFormation, Docker, Go, Kafka, Kubernetes.
What seniority level is this role?
Gamma targets senior candidates for this position.
Is this full-time or contract?
Full-time for this Site Reliability Engineer role at Gamma.

From the original posting

About the role

Gamma's infrastructure needs to be rock-solid for millions of daily users while enabling our engineering teams to ship fast. You'll own the operational health of our full backend platform, building automation and tooling that improves reliability and partnering with engineering to design systems that are observable, resilient, and easy to operate. Your work directly impacts every Gamma user's experience.

This is a high-impact role where you'll balance reliability with velocity, knowing when to move fast and when to prioritize stability. You'll lead incident response, drive systemic improvements, and help shape how Gamma scales to serve its next 100 million users.

Our team has a strong in-office culture and works in person 4–5 days per week in San Francisco. We love working together to stay creative and connected, with flexibility to work from home when focus matters most.

What you'll do

  • Own the reliability, availability, and performance of Gamma's production systems across our AWS infrastructure

  • Build observability infrastructure from the ground up: metrics, logging, tracing, and alerting that give the team genuine visibility into system health before users feel the impact

  • Design and ship automation that reduces toil, makes deployments safer, and gets us back on our feet faster when things go wrong

  • Lead incident response and blameless post-mortems, then follow through on the systemic fixes that keep the same issues from coming back

  • Partner with engineering teams on architecture reviews, SLO and SLI design, and reliability best practices that scale with the product

  • Manage and optimize our compute, networking, databases, and managed services

What you'll bring

  • 5+ years in site reliability engineering, DevOps, or systems engineering with deep, hands-on AWS expertise

  • Strong programming skills in Python, Go, or TypeScript/Node.js, applied to building real tools and automation

  • Solid experience with infrastructure-as-code (Terraform, CloudFormation) and end-to-end observability solutions

  • Track record of making systems meaningfully more reliable through automation, smarter monitoring, and architectural improvements

  • Deep understanding of networking, distributed systems, containerization (Docker, Kubernetes), and database performance at scale

  • Sharp incident management instincts and the debugging skills to navigate complex production failures

  • Experience scaling SaaS products to millions of users, or background with Kafka, chaos engineering, or service mesh technologies (Nice to have)

  • AWS certifications, or experience with security and compliance frameworks like SOC 2 or ISO 27001 (Nice to have)

Compensation range:

The base salary for this full-time position, which spans multiple internal levels depending on qualifications, ranges between $230K - $310K plus benefits & equity.

Final offer amounts are determined by multiple factors, including but not limited to experience and expertise in the requirements listed above.

If you're interested in this role but you don't meet every requirement, we encourage you to apply anyway! We're always excited about meeting great people.

Similar positions

Gamma
Software Engineer, Platform
Gamma · San Francisco
Picogrid
Site Reliability Engineer
Picogrid · El Segundo, CA
Turquoise Health
Platform Operations Engineer
Turquoise Health · Remote
Browserbase
Software Engineer (Core Infrastructure)
Browserbase · San Francisco