← Back to results

Member of Technical Staff - Kernels & GPU Performance

Build and optimize low-level execution primitives for AI inference across various hardware architectures at Gimlet Labs.

Location
San Francisco, CA, United States
Compensation
Not disclosed
Level
staff
Type
full time

Posted by employer 6 months ago

First seen on Joblaze 5 hours ago

Last verified on the company career page 5 hours ago

Apply at Gimlet Labs → Save job Scanned from gimletlabs.ai

Skills & Technologies

What you'll build

  • Build and optimize kernels for production AI workloads
  • Develop execution strategies for accelerator architectures
  • Improve memory efficiency and scheduling behavior
  • Partner with engineers for performance optimization
  • Establish performance engineering standards

Must have

  • Strong software engineering fundamentals
  • Experience working on performance-critical systems
  • Bachelor's degree in a relevant field

Nice to have

  • Experience with CUDA, Triton, CUTLASS
  • Deep understanding of GPU execution models
  • Experience optimizing memory access patterns
  • Familiarity with occupancy and latency hiding
  • Experience using profiling and performance analysis tools

Requirements

Education
Bachelor's degree

Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.

Joblaze summary

In this role, the Member of Technical Staff focuses on developing and optimizing low-level execution primitives to enhance AI inference performance across various hardware architectures. Key skills include a strong foundation in software engineering and experience with performance-critical systems, particularly in relation to GPU execution models and memory hierarchies. This position is well-suited for individuals with a background in performance optimization and a deep understanding of accelerator programming. Gimlet Labs is in a growth phase, expanding its technology into a production neocloud, which presents opportunities to tackle complex challenges.

Joblaze insights

  • Listed today — first seen on Joblaze October 5, 2026. Last confirmed on Gimlet Labs's careers page October 5, 2026.
  • CUDA appears in 4.8% of 269 comparable staff ai/ml roles in United States; CUTLASS appears in 0.7% of 269 comparable staff ai/ml roles in United States.

Quick facts

What's the tech stack?
Joblaze extracted these technologies from the posting: CUDA, CUTLASS, GPU, Triton.
What seniority level is this role?
Gimlet Labs targets staff-level candidates for this position.
Is this full-time or contract?
Full-time for this Member of Technical Staff - Kernels & GPU Performance role at Gimlet Labs.

From the original posting

About the role

As a Member of Technical Staff, you will build and optimize the low-level execution primitives that turn accelerator performance into production inference performance.

Rather than optimizing for one hardware architecture, you will work across accelerators with different execution models, memory hierarchies, capabilities, and software stacks. Your work will shape the latency, throughput, and efficiency Gimlet can achieve across established and emerging hardware architectures.

You will work close to the hardware across kernel implementation, memory access, execution behavior, profiling, and performance validation. You will develop optimizations that account for differences between accelerator architectures and partner with compiler, ML systems, and distributed systems engineers to improve performance across the full execution stack.

What success looks like

In the first 12-18 months, you will:

  • Build and optimize kernels that improve latency, throughput, and hardware utilization for production AI workloads

  • Develop execution strategies that unlock performance across both established and emerging accelerator architectures

  • Improve memory efficiency, scheduling behavior, and execution characteristics across the inference stack

  • Partner with compiler, runtime, and distributed systems engineers to ensure end-to-end performance optimization

  • Influence how heterogeneous hardware is deployed and utilized within the next generation of AI infrastructure

  • Help establish performance engineering standards that shape the future of Gimlet's execution platform

  • Strong software engineering fundamentals

  • Experience working on performance-critical systems close to hardware

  • Comfort reasoning about low-level execution behavior, memory hierarchies, and performance tradeoffs

  • Bachelor's degree in a relevant field, or an equivalent combination of education, training, and professional experience.

Strong candidates may also have

  • Experience with CUDA, Triton, CUTLASS, or other accelerator programming models

  • Deep understanding of GPU execution models (warps/wavefronts, blocks, grids)

  • Experience optimizing memory access patterns (coalescing, shared memory, cache behavior)

  • Familiarity with occupancy, latency hiding, and instruction-level parallelism

  • Experience using profiling and performance analysis tools

  • Familiarity with multi-GPU or distributed execution is a plus

  • Solve hard problems.

  • Own meaningful work.

  • Build for production.

  • Help define what’s next.

Standard company text repeated across Gimlet Labs's postings is omitted here.

Similar positions

Gimlet Labs
Member of Technical Staff - Distributed Systems
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - ML Systems & Inference
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - Infrastructure
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Member of Technical Staff - Compiler Engineer
Gimlet Labs · San Francisco, CA, United States
Gimlet Labs
Network Engineer
Gimlet Labs · San Francisco, CA, United States