"tensorrt llm" Jobs

56 open tech roles matching “tensorrt llm”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, vLLM, TensorRT-LLM. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 56 results

Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.

Inferact Remote Published 1 week ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Hippocratic AI

Own the serving infrastructure for healthcare AI, optimizing LLM inference systems to enhance patient experiences.

Hippocratic AI Menlo Park, CA Published 2 weeks ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focused on ML performance to optimize large language models in a fast-paced startup environment.

baseten San Francisco Published 2 years ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco, California, United States Published 1 day ago
Flexible on stack
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Periodic Labs

Join Periodic Labs as an ML Systems Engineer to build and optimize large-scale training and reinforcement learning infrastructure.

Periodic Labs Menlo Park, CA $250k–$350k/yr Published 4 months ago
Flexible on stack
Ambient

Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced machine learning models.

Ambient Redwood City, United States Published 1 day ago
Flexible on stack 70% coding
Ambient

Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced AI models.

Ambient Redwood City Published 1 month ago
Flexible on stack 70% coding
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings