← Back to results

Member of Technical Staff - ML Systems & Inference

Build inference systems for AI models in production at Gimlet Labs, focusing on performance and efficiency.

Location
San Francisco, CA, United States
Compensation
Not disclosed
Level
staff
Type
full time

Posted by employer 6 months ago

First seen on Joblaze 5 hours ago

Last verified on the company career page 5 hours ago

Apply at Gimlet Labs → Save job Scanned from gimletlabs.ai

Skills & Technologies

What you'll build

  • Build inference systems that execute models end-to-end in production
  • Determine how inference executes across the pipeline
  • Improve latency, throughput, and efficiency of production inference workloads
  • Design execution strategies across batching, scheduling, concurrency, and resource utilization
  • Enable new models and inference techniques to run efficiently in production

Must have

  • Strong software engineering fundamentals
  • Experience building or operating ML inference or model serving systems
  • Bachelor's degree in a relevant field

Nice to have

  • Experience with inference runtimes such as TensorRT-LLM, vLLM
  • Deep understanding of modern model architectures and attention mechanisms
  • Experience with batching, scheduling, and concurrency control in inference systems
  • Familiarity with KV cache management and memory placement strategies
  • Experience profiling and tuning latency- and throughput-critical systems

Requirements

Education
Bachelor's degree

Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.

Joblaze summary

In this role, the Member of Technical Staff will focus on developing and optimizing machine learning inference systems for production environments. Key responsibilities include managing request batching, scheduling, and memory placement while ensuring efficient execution across various hardware setups. This position is well-suited for candidates with a strong software engineering background and experience in ML inference systems, particularly those familiar with performance tuning and modern model architectures. Gimlet Labs is currently expanding its technology offerings, providing a unique environment for tackling complex challenges.

Joblaze insights

  • Listed today — first seen on Joblaze October 5, 2026. Last confirmed on Gimlet Labs's careers page October 5, 2026.
  • Python appears in 51.7% of 269 comparable staff ai/ml roles in United States; TensorRT-LLM appears in 1.5% of 269 comparable staff ai/ml roles in United States.

Quick facts

What's the tech stack?
Joblaze extracted these technologies from the posting: C++, Python, TensorRT-LLM, vLLM.
What seniority level is this role?
Gimlet Labs targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Member of Technical Staff - ML Systems & Inference role at Gimlet Labs.

From the original posting

About the role

As a Member of Technical Staff focused on ML Systems, you will build the inference systems that execute models end-to-end in production.

You will work on the systems that determine how inference executes across that pipeline: how requests are batched and scheduled, how stages are placed and scaled, how KV cache and intermediate state move between accelerators, and how the system balances latency, throughput, and utilization across different hardware characteristics.

You will work across model serving, batching, scheduling, concurrency, KV cache management, and memory placement. You will help bring up models on novel hardware. You will support new model architectures and inference techniques, improve performance under real production workloads, and partner with compiler, kernel, networking, and distributed systems engineers to optimize the full execution path.

What success looks like

In the first 12-18 months, you will:

  • Improve the latency, throughput, and efficiency of production inference workloads

  • Design execution strategies across batching, scheduling, concurrency, and resource utilization

  • Improve KV cache management, memory efficiency, and execution under load

  • Enable new models, accelerator architectures, and inference techniques to run efficiently in production

  • Strong software engineering fundamentals

  • Experience building or operating ML inference or model serving systems

  • Comfort reasoning about performance, memory usage, and system behavior under load

  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have

  • Experience with inference runtimes such as TensorRT-LLM, vLLM, or custom serving systems

  • Deep understanding of modern model architectures and attention mechanisms

  • Experience with batching, scheduling, and concurrency control in inference systems

  • Familiarity with KV cache management and memory placement strategies

  • Experience profiling and tuning latency- and throughput-critical systems

  • Software development experience in Python and C++

  • Solve hard problems.

  • Own meaningful work.

  • Build for production.

  • Help define what’s next.

Standard company text repeated across Gimlet Labs's postings is omitted here.

Similar positions

Gimlet Labs
Member of Technical Staff - Distributed Systems
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - Kernels & GPU Performance
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - Infrastructure
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - Compiler Engineer
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Network Engineer
Gimlet Labs · San Francisco, CA, United States