Direct from source · No middlemen

Staff Vllm Jobs

16 open positions · Updated 1 week ago

Average salary (USD/year): 196.4k–329.9k/yr
16 roles Staff 16 5–12 yrs experience New York +5 more Full-time newest 1 week ago

16 Vllm roles across 6 companies, most in AI/ML; 13% fully remote; typical advertised salary $192k

Work arrangement: 2 fully remote · 4 hybrid · 3 on-site · 7 not stated

Advertised salaries: p25 $188k · median $192k · p75 $200k (from 13 disclosed annual salaries, USD)

Counts are open roles Joblaze currently tracks on company career pages for this exact skill/location; salaries are advertised minimums, annual, converted to USD.

Showing 16 of 16 positions

Search with filters →
Inferact

Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.

Inferact Remote Published 1 week ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Stuut

Join Stuut as a Member of the Technical Staff to design and deploy AI-powered systems for financial operations.

Stuut San Francisco Published 1 month ago
Flexible on stack
Coreweave

Join CoreWeave as a Staff Software Engineer to lead the development of a Kubernetes-native inference platform for AI workloads.

Coreweave Sunnyvale, CA / Bellevue, WA $188k–$275k/yr Published 4 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding