← Back to results

Member of Technical Staff - Infrastructure

Build reliable production infrastructure for Gimlet's AI cloud as an Infrastructure Platform Engineer.

Location
San Francisco, CA, United States
Compensation
Not disclosed
Level
staff
Type
full time

Posted by employer 3 months ago

First seen on Joblaze 1 day ago

Last verified on the company career page 1 day ago

Apply at Gimlet Labs → Save job Scanned from gimletlabs.ai

What you'll build

  • Deploy and operate production clusters across different accelerator architectures
  • Automate hardware provisioning, validation, upgrades, and fleet lifecycle management
  • Improve cluster scheduling, resource utilization, isolation, and capacity management
  • Build observable infrastructure for faster debugging and incident response
  • Partner across teams to bring new accelerators into production

Must have

  • Experience in infrastructure, cluster engineering, platform engineering, SRE, or HPC
  • Strong Linux systems knowledge and production debugging experience
  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems
  • Experience automating infrastructure with Python, Go, Terraform, Ansible, or similar tools
  • Experience with GPU or accelerator infrastructure

Nice to have

  • Experience building or operating AI inference, training, HPC, or neocloud infrastructure
  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines
  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling
  • Experience debugging distributed workload performance
  • Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana

Requirements

Education
Bachelor's degree

Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.

Joblaze summary

In the role of Infrastructure Platform Engineer at Gimlet Labs, the individual will focus on developing systems that transform diverse accelerator hardware into dependable production infrastructure for AI applications. Key skills include expertise in Linux, Kubernetes, and automation tools like Python and Terraform, alongside experience with GPU infrastructure. This position is well-suited for candidates with a background in infrastructure or platform engineering, particularly those familiar with high-performance computing environments. Gimlet is currently expanding its technology offerings, presenting opportunities to tackle complex challenges in a growing company.

Joblaze insights

  • Listed yesterday — first seen on Joblaze October 5, 2026. Last confirmed on Gimlet Labs's careers page October 5, 2026.
  • Kubernetes appears in 63% of 81 comparable staff devops/sre roles in United States; OpenTelemetry appears in 3.7% of 81 comparable staff devops/sre roles in United States.

Quick facts

What's the tech stack?
Joblaze extracted these technologies from the posting: Ansible, CUDA, Go, Grafana, Kubernetes, Linux.
What seniority level is this role?
Gimlet Labs targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Member of Technical Staff - Infrastructure role at Gimlet Labs.

From the original posting

About the role

As an Infrastructure Platform Engineer, you will build the systems that turn heterogeneous accelerator hardware into reliable production infrastructure for Gimlet's AI cloud.

Gimlet's fleet spans hardware with different architectures, software stacks, operational characteristics, and failure modes. Your work will determine how new hardware is brought online, how clusters are provisioned and operated, and how production inference systems remain reliable as the fleet scales.

You will work across bare metal, Linux, Kubernetes, cluster scheduling, observability, and automation. You will build systems that abstract differences between accelerator architectures, make new hardware production-ready, and improve the reliability and operability of Gimlet’s infrastructure.

What success looks like

In your first 12–18 months, you will:

  • Deploy and operate production clusters across different accelerator architectures

  • Automate hardware provisioning, validation, upgrades, and fleet lifecycle management

  • Improve cluster scheduling, resource utilization, isolation, and capacity management

  • Build observable infrastructure that enables faster debugging, incident response, and recovery

  • Partner across distributed systems, runtime, compiler, networking, and hardware teams to bring new accelerators into production

  • Experience in infrastructure, cluster engineering, platform engineering, SRE, or HPC

  • Strong Linux systems knowledge and production debugging experience

  • Experience operating Kubernetes, Slurm, Nomad, or similar orchestration systems

  • Experience automating infrastructure with Python, Go, Terraform, Ansible, or similar tools

  • Experience with GPU or accelerator infrastructure, including drivers, firmware, or CUDA/ROCm

  • The ability to build systems that are observable, recoverable, and reliable in production

  • A bachelor’s degree in a relevant field or equivalent practical experience

Strong candidates may also have

  • Experience building or operating AI inference, training, HPC, or neocloud infrastructure

  • Experience with bare-metal provisioning, PXE/iPXE, image pipelines, BIOS/firmware management, or rack bring-up

  • Experience with multi-tenant cluster isolation, quota systems, fair scheduling, or usage accounting

  • Experience debugging distributed workload performance across compute, memory, network, and storage bottlenecks

  • Experience building observability platforms using technologies such as Prometheus, OpenTelemetry, Grafana, or similar tooling

  • Familiarity with heterogeneous hardware environments across NVIDIA, AMD, Intel, ARM, or emerging accelerators

  • Solve hard problems.

  • Own meaningful work.

  • Build for production.

  • Help define what’s next.

Standard company text repeated across Gimlet Labs's postings is omitted here.

Similar positions

Gimlet Labs
Member of Technical Staff - Distributed Systems
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - Kernels & GPU Performance
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Network Engineer
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - ML Systems & Inference
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Data Center Facilities Operations Lead
Gimlet Labs · San Francisco, CA, United States