"llm inference" Jobs
438 open tech roles matching “llm inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 438 results
Own the serving infrastructure for healthcare AI, optimizing LLM inference systems to enhance patient experiences.
Join Ricursive Intelligence to tackle challenges in scaling and optimization for LLM training and inference.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.
Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.
Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.
Join Inferact as the first in-house legal hire to lead legal functions and support a fast-growing AI inference company.
Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.
Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.
Own the inference systems that power frontier AI models in production and research at a tech-first startup.
Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.
Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.