"inference runtimes" Jobs

41 open tech roles matching “inference runtimes”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 41 results

Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 4 days ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
baseten

Lead the Runtime Fabric team at Baseten to build container runtimes tailored for AI inference workloads.

baseten San Francisco Published 3 months ago
Flexible on stack Heavy meetings
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a lead Software Engineer to own and develop production-grade Voice AI systems that impact daily lives.

baseten San Francisco Published 4 months ago
Flexible on stack
baseten

Join Baseten as a Site Reliability Engineer to enhance the reliability of our multi-cloud Kubernetes infrastructure.

baseten San Francisco Published 4 months ago
Flexible on stack 70% coding
Coreweave

Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA $207k–$275k/yr Published 2 months ago
Flexible on stack