Direct from source · No middlemen

Vllm Jobs

52 open positions · Updated 1 week ago

Average salary (USD/year): 184.8k–290.2k/yr
52 roles Senior 22, Staff 16, Mid 12, Junior 1, Lead 1 2–12 yrs experience Sunnyvale, CA / Bellevue, WA +26 more Full-time newest 1 week ago

52 Vllm roles across 24 companies, most in AI/ML; 10% fully remote; typical advertised salary $190k

Work arrangement: 5 fully remote · 11 hybrid · 11 on-site · 25 not stated

Advertised salaries: p25 $157.8k · median $190k · p75 $200k (from 25 disclosed annual salaries, USD)

Counts are open roles Joblaze currently tracks on company career pages for this exact skill/location; salaries are advertised minimums, annual, converted to USD.

Showing 20 of 52 positions

Search with filters →
Inferact

Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.

Inferact Remote Published 1 week ago
Flexible on stack
ChipAgents

Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.

ChipAgents San Jose $150k–$350k/yr Published 3 months ago
Flexible on stack
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Sarvam AI

Own the model lifecycle for defence and strategic sector deployments as an MLOps Engineer at Sarvam AI.

Sarvam AI Delhi Published 4 months ago
Flexible on stack
Sarvam AI

Own the full lifecycle of AI system deployments as a Strategic Deployment Engineer at Sarvam, working directly with clients in complex environments.

Sarvam AI Delhi Published 4 months ago
Flexible on stack
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Handshake

Join Handshake as a Senior Software Engineer to build scalable ML infrastructure for a fast-growing AI data business.

Handshake San Francisco, CA Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
baseten

Join Baseten as a Solutions Architect to translate business needs into technical solutions for AI deployments.

baseten San Francisco Published 6 months ago
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Dialpad

Join Dialpad as a Software Engineer to build and improve ML inference systems for AI models at scale.

Dialpad Buenos Aires, Argentina Published 2 months ago
Flexible on stack 70% coding
Page 1 of 3 Next →