← Back to results

Staff Site Reliability Engineer

Join Hippocratic AI as a Staff Site Reliability Engineer to manage a fleet of GPU-backed models and enhance healthcare AI systems.

Location
Menlo Park, CA, United States
Compensation
Not disclosed
Level
staff
Type
full time

Posted by employer 1 day ago

First seen on Joblaze 12 hours ago

Last verified on the company career page 12 hours ago

Apply at Hippocratic AI → Save job Scanned from hippocraticai.com

What you'll build

  • Design and build GPU management and scheduling platform
  • Build metrics pipeline for GPU load and utilization data
  • Implement admission control for capacity protection
  • Build autoscaling for model replicas
  • Develop cloud orchestration systems in Python and Go

Must have

  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering
  • Computer Science Degree Required from a top CS program
  • Strong software engineering fundamentals in Python and/or Go
  • Experience designing systems that make decisions from operational metrics
  • Deep experience with infrastructure automation and CI/CD
  • Hands-on production experience with at least one major cloud platform

Nice to have

  • Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators
  • Familiarity with ML inference serving and model deployment
  • Experience with Kubernetes autoscaling internals
  • Experience implementing HIPAA and SOC 2 compliance
  • Experience operating in an HPC environment
  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field

Requirements

Experience
10+ years
Education
Bachelor's degree

Not disclosed in this posting: compensation, work arrangement, visa sponsorship.

Joblaze summary

In this role, the Staff Site Reliability Engineer at Hippocratic AI is responsible for designing and building a GPU management and scheduling platform that optimizes the performance of approximately 30 models across diverse hardware. The position requires strong software engineering skills in Python and Go, along with extensive experience in site reliability and DevOps practices, particularly in cloud environments. Ideal candidates will have over a decade of experience and a solid background in systems engineering, making them well-suited to tackle complex infrastructure challenges. The role also involves mentoring team members and collaborating with engineers and researchers to enhance oper

Joblaze insights

  • Listed today — first seen on Joblaze September 19, 2026. Last confirmed on Hippocratic AI's careers page September 19, 2026.
  • Kubernetes appears in 66.7% of 78 comparable staff devops/sre roles in United States; Grafana appears in 10.3% of 78 comparable staff devops/sre roles in United States.

Quick facts

How much experience is required?
At least 10 years of relevant experience for this Staff Site Reliability Engineer role.
What's the tech stack?
Joblaze extracted these technologies from the posting: AWS, Azure, CI/CD, Datadog, Docker, ELK.
What seniority level is this role?
Hippocratic AI targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Staff Site Reliability Engineer role at Hippocratic AI.

From the original posting

About the Role

We're looking for a Senior Site Reliability Engineer who is equally at home writing production software and running the infrastructure it lives on — and who wants to take ownership of one of the hardest, highest-leverage problems on our platform: intelligently managing a large fleet of GPU-backed models.

We run nearly 30 models across heterogeneous hardware, and keeping that fleet fast, reliable, and cost-effective is a serious engineering challenge. You'll build the GPU management and scheduling platform that sits at the center of it — collecting utilization and load metrics, interpreting what they actually mean, and using them to make real-time decisions about admission control and scaling. The goal: route and schedule inference calls so we use our capacity efficiently without exceeding it, and scale model replicas up and down automatically as demand shifts.

This is a senior role for someone with a decade in the field who can move fluidly between systems engineering and software development, and who is excited to own a complex, evolving system end to end.

What You'll Do

  • Design and build our GPU management and scheduling platform — the system that decides when, where, and how inference calls run across a fleet of ~30 models on heterogeneous hardware

  • Build the metrics pipeline that collects GPU load and utilization data, and the logic that turns those signals into decisions

  • Implement admission control to protect capacity — deciding when to accept, queue, or shed inference requests so we operate within fleet limits

  • Build autoscaling that adjusts the number of model replicas in response to real-time demand and utilization

  • Develop cloud orchestration systems and operators in Python and Go to manage the model fleet

  • Architect and operate scalable, fault-tolerant, secure production systems on AWS, GCP, or Azure

  • Design and build infrastructure automation and deployment pipelines (Terraform, CI/CD) as first-class software

  • Stand up and maintain monitoring, logging, and alerting that keep the platform reliable and performant

  • Develop and enforce security and compliance policies appropriate to a healthcare AI platform

  • Partner with engineers and research scientists to diagnose and resolve complex infrastructure, deployment, and operational issues

  • Mentor engineers and raise the technical bar across the team

What You Bring

Must-Have

  • 10+ years of professional experience across site reliability / DevOps engineering and software engineering

  • Computer Science Degree Required from a top CS program.

  • Strong software engineering fundamentals — you build orchestration and scheduling systems in Python and/or Go, not just configure off-the-shelf tools

  • Experience designing systems that make decisions from operational metrics — collecting signals, interpreting them, and driving control loops such as autoscaling, load shedding, or admission control

  • Deep experience with infrastructure automation and CI/CD (Terraform, GitLab CI/CD, or similar)

  • Hands-on production experience with at least one major cloud platform (AWS, GCP, or Azure)

  • Strong knowledge of containerization and orchestration (Docker, Kubernetes)

  • Experience with monitoring and logging stacks (ELK, Grafana, Datadog, or similar)

  • Familiarity with secrets management and security tooling (HashiCorp Vault, AWS KMS, Azure Key Vault)

  • Excellent problem-solving skills and the ability to work both independently and collaboratively

  • Strong communication and interpersonal skills

Nice-to-Have

  • Experience managing GPU fleets or scheduling workloads across heterogeneous accelerators

  • Familiarity with ML inference serving and model deployment (e.g. Triton, KServe, Ray Serve, or similar)

  • Experience with Kubernetes autoscaling internals (HPA/VPA, custom metrics, custom controllers)

  • Experience implementing HIPAA and SOC 2 compliance

  • Experience operating in an HPC environment

  • Bachelor's or Master's in Computer Science, Computer Engineering, or a related field

Join our team at Hippocratic AI and help shape the future of clinically safe, production-grade AI systems.

Standard company text repeated across Hippocratic AI's postings is omitted here.

Similar positions

Fal
Software Engineer, Platform
Fal · San Francisco
baseten
Capacity Operations Manager
baseten · San Francisco
Fal
Software Engineer, Platform
Fal · Remote - Global
Qualified Health
Senior DevOps Engineer
Qualified Health · Palo Alto - Hybrid
HappyRobot
Site Reliability Engineer
HappyRobot · San Francisco