"tensorrt llm" Jobs

56 open tech roles matching “tensorrt llm”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, vLLM, TensorRT-LLM. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 56 results

baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Together AI

Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.

Together AI Singapore Published 1 month ago
Flexible on stack 70% coding
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems for generative AI with leading organizations.

Fireworks AI Singapore Published 1 month ago
Flexible on stack 70% coding
DeepL

Lead research on fine-tuning and steerability of LLM-based translation models in a collaborative AI-focused environment.

DeepL London Published 1 month ago
Flexible on stack
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as an AI Field Engineer to build production systems for generative AI with large organizations across EMEA.

Fireworks AI London Published 1 month ago
Flexible on stack 70% coding
Twelve Labs

Lead and build a new team focused on developing Jockey Core, a reasoning LLM for video understanding at Twelve Labs.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
Cloudflare

Join Cloudflare as a Senior Machine Learning Engineer to optimize and productionize ML models for a global serverless inference platform.

Cloudflare Hybrid Published 2 months ago
Flexible on stack
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $92k–$135k/yr Published 10 months ago
Illumio

Architect high-scale distributed systems and lead the development of autonomous AI agents in a dynamic cybersecurity environment.

Illumio HQ - Sunnyvale (Office) Published 2 months ago
Flexible on stack
Sarvam AI

Own the architecture of Sarvam's vision models serving harness, ensuring high-quality document intelligence at national scale.

Sarvam AI Bengaluru Published 3 weeks ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems and engage with enterprise customers on generative AI solutions.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems for innovative AI-native companies.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding
Roboflow

Join Roboflow as a Machine Learning Engineer to enhance our inference engine and contribute to impactful computer vision projects.

Roboflow NY, SF or Remote Published 2 months ago
Flexible on stack