Build systems for scheduling and coordinating AI workloads across Gimlet’s infrastructure as a Member of Technical Staff.
Posted by employer 7 months ago
First seen on Joblaze 1 day ago
Last verified on the company career page 1 day ago
Skills & Technologies
What you'll build
Must have
Nice to have
AI in the day-to-day
We combine large-scale compute infrastructure with an execution platform that partitions AI workloads.
Requirements
Not disclosed in this posting: compensation, work arrangement, visa sponsorship.
Joblaze summary
In the role of Member of Technical Staff at Gimlet Labs, the individual will focus on developing systems that manage the scheduling and coordination of AI workloads across a diverse infrastructure. Key skills include experience with distributed systems, strong software engineering fundamentals, and familiarity with resource management and scheduling techniques. This position is well-suited for candidates with a background in systems development and a solid understanding of concurrency and fault tolerance. Gimlet is in a growth phase, expanding its technology to support new hardware and data centers.
Joblaze insights
Quick facts
From the original posting
About the role
As a Member of Technical Staff, you will build the systems that schedule, route, and coordinate AI workloads across Gimlet’s infrastructure.
Different stages of an inference pipeline may run on different hardware, scale independently, and exchange state across the system. Your work will determine how those workloads are placed, coordinated, routed, recovered, and operated in production.
You will work across scheduling, orchestration, control planes, APIs, and fault tolerance. You will design systems that make distributed infrastructure easier to operate, enable workloads to run reliably across a heterogeneous fleet, and partner with compiler, ML systems, networking, and infrastructure engineers to connect the full execution stack.
What success looks like
In the first 12-18 months, you will:
Build scheduling and orchestration systems for heterogeneous compute
Design systems that manage independently scalable stages of distributed inference pipelines
Improve the reliability and fault tolerance of production AI infrastructure
Develop control planes and APIs that simplify how workloads are deployed and managed
Improve resource management and scheduling as Gimlet expands across new accelerator types, nodes, and data centers
Help the platform scale across additional hardware, nodes, and data centers
Experience building or operating distributed systems in production
Strong software-engineering and systems fundamentals
The ability to reason about concurrency, consistency, failure modes, and system tradeoffs
Experience with scheduling, resource management, RPC, or asynchronous messaging
A bachelor’s degree in a relevant field or equivalent practical experience
Strong candidates may also have
Experience with Kubernetes or Kubernetes-adjacent systems beyond basic usage
Experience designing service-oriented architectures using RPC or asynchronous messaging
Familiarity with scheduling, queues, or resource management systems
Experience building reliable APIs and operating systems under high load
Software development experience in languages commonly used for systems development (e.g., Go, C++, Python)
Solve hard problems.
Own meaningful work.
Build for production.
Help define what’s next.
Standard company text repeated across Gimlet Labs's postings is omitted here.