← Back to jobs

vLLM

52 jobs tagged vLLM

Median salary: $190,000 (from 25 disclosed salaries, annual, USD) 10% remote Top hiring: Inferact, Together AI, Coreweave, Fireworks AI, baseten

vLLM jobs — quick facts

How many vLLM jobs are available?
52 active vLLM roles sourced directly from company career pages, updated daily on Joblaze.
What is the median salary for vLLM jobs?
The median advertised salary for vLLM roles on Joblaze is about $190,000 per year.
Which companies are hiring for vLLM?
Companies currently hiring vLLM include Inferact, Together AI, Coreweave, Fireworks AI, baseten.
Are vLLM jobs remote?
10% of active vLLM roles on Joblaze offer remote work.
Inferact

Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.

Inferact Remote Published 1 week ago
Flexible on stack
ChipAgents

Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.

ChipAgents San Jose $150k–$350k/yr Published 3 months ago
Flexible on stack
Sarvam AI

Own the full lifecycle of AI system deployments as a Strategic Deployment Engineer at Sarvam, working directly with clients in complex environments.

Sarvam AI Delhi Published 4 months ago
Flexible on stack
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Sarvam AI

Own the model lifecycle for defence and strategic sector deployments as an MLOps Engineer at Sarvam AI.

Sarvam AI Delhi Published 4 months ago
Flexible on stack
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Handshake

Join Handshake as a Senior Software Engineer to build scalable ML infrastructure for a fast-growing AI data business.

Handshake San Francisco, CA Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Solutions Architect to translate business needs into technical solutions for AI deployments.

baseten San Francisco Published 6 months ago
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Dialpad

Join Dialpad as a Software Engineer to build and improve ML inference systems for AI models at scale.

Dialpad Buenos Aires, Argentina Published 2 months ago
Flexible on stack 70% coding
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Sola

Join Sola as a Software Engineer to build reliable AI features in a fast-growing startup environment in NYC.

Sola New York City Published 5 months ago
Flexible on stack
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 1 day ago
Flexible on stack
Stuut

Join Stuut as a Member of the Technical Staff to design and deploy AI-powered systems for financial operations.

Stuut San Francisco Published 1 month ago
Flexible on stack
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Coreweave

Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA $207k–$275k/yr Published 2 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Staff Software Engineer to lead the development of a Kubernetes-native inference platform for AI workloads.

Coreweave Sunnyvale, CA / Bellevue, WA $188k–$275k/yr Published 4 months ago
Flexible on stack
Coreweave

Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.

Coreweave Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA $188k–$275k/yr Published 5 months ago
Flexible on stack
ElevenLabs

Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.

ElevenLabs United Kingdom Published 2 weeks ago
Flexible on stack
SpaceX

Join SpaceX as a Software Engineer to develop high-performance AI inference systems for mission-critical applications.

SpaceX Palo Alto, CA $135k–$175k/yr Published 3 weeks ago
Flexible on stack
Together AI

Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.

Together AI Singapore Published 1 month ago
Flexible on stack 70% coding
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems for generative AI with leading organizations.

Fireworks AI Singapore Published 1 month ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as an AI Field Engineer to build production systems for generative AI with large organizations across EMEA.

Fireworks AI London Published 1 month ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems and engage with enterprise customers on generative AI solutions.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems for innovative AI-native companies.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding
Cloudflare

Join Cloudflare as a Senior Machine Learning Engineer to optimize and productionize ML models for a global serverless inference platform.

Cloudflare Hybrid Published 2 months ago
Flexible on stack
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
Roboflow

Join Roboflow as a Machine Learning Engineer to enhance our inference engine and contribute to impactful computer vision projects.

Roboflow NY, SF or Remote Published 2 months ago
Flexible on stack
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $92k–$135k/yr Published 10 months ago
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $139k–$204k/yr Published 7 months ago
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $165k–$242k/yr Published 11 months ago