"vllm" Jobs

53 open tech roles matching “vllm”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, vLLM, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 53 results

Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 3 days ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Site Reliability Engineer to enhance the reliability and performance of AI inference systems at scale.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack AI-first team
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Inferact

Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Inferact

Join Inferact as a Product Marketing Manager to enhance vLLM's presence in the AI inference space through strategic marketing and community engagement.

Inferact San Francisco Published 1 month ago
Inferact

Join Inferact as a Founding Product Designer to shape the visual identity and user experience of our AI inference engine.

Inferact San Francisco Published 1 month ago
Flexible on stack
Inferact

Join Inferact as an IT Support & Operations Engineer to enhance internal technology and security for a growing AI startup.

Inferact San Francisco $125k–$170k/yr Published 3 days ago
Inferact

Lead HR and People Operations at Inferact, scaling infrastructure in a fast-paced startup environment.

Inferact San Francisco $180k–$250k/yr Published 4 days ago
Inferact
Head of Legal Hybrid Visa

Join Inferact as the first in-house legal hire to lead legal functions and support a fast-growing AI inference company.

Inferact San Francisco Published 4 days ago
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack