Direct from source · No middlemen

Vllm Jobs in San Francisco

24 open positions · Updated 1 week ago

Average salary (USD/year): 205.8k–311k/yr
24 roles Senior 9, Staff 9, Mid 5, Lead 1 2–10 yrs experience *HQ - San Francisco, CA +8 more Full-time newest 1 week ago

24 Vllm roles across 12 companies, most in AI/ML; typical advertised salary $200k

Work arrangement: 6 hybrid · 1 on-site · 17 not stated

Advertised salaries: p25 $191k · median $200k · p75 $203.5k (from 15 disclosed annual salaries, USD)

Counts are open roles Joblaze currently tracks on company career pages for this exact skill/location; salaries are advertised minimums, annual, converted to USD.

Showing 20 of 24 positions

Search with filters →
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Handshake

Join Handshake as a Senior Software Engineer to build scalable ML infrastructure for a fast-growing AI data business.

Handshake San Francisco, CA Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Solutions Architect to translate business needs into technical solutions for AI deployments.

baseten San Francisco Published 6 months ago
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 2 days ago
Flexible on stack
Stuut

Join Stuut as a Member of the Technical Staff to design and deploy AI-powered systems for financial operations.

Stuut San Francisco Published 1 month ago
Flexible on stack
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Coreweave

Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA $207k–$275k/yr Published 2 months ago
Flexible on stack
Coreweave

Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.

Coreweave Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA $188k–$275k/yr Published 6 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
Page 1 of 2 Next →