"inference optimization" Jobs

588 open tech roles matching “inference optimization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 588 results

Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft Seattle, WA $136.2k–$170.2k/yr Published 3 months ago
Flexible on stack
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft New York, NY $148k–$185k/yr Published 3 months ago
Flexible on stack
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $139k–$204k/yr Published 7 months ago
Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft San Francisco, CA $148k–$185k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale in a fully remote role.

Inferact Remote Published 2 weeks ago
Flexible on stack
baseten

Join Baseten as an AI Inference Engineer to architect and deploy high-scale production AI applications while collaborating with customers.

baseten San Francisco Published 1 month ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as a Software Engineer focused on Performance Optimization to enhance AI infrastructure efficiency and speed.

Fireworks AI San Mateo Published 1 year ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Volta

Lead product strategy for inference infrastructure and token-serving capabilities in a rapidly growing AI infrastructure company.

Volta Palo Alto, CA Published 1 month ago
baseten

Join Baseten as a GPU Kernel Engineer to optimize high-performance GPU kernels for cutting-edge AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack 70% coding
Iambic Therapeutics

Join Iambic Therapeutics as a Machine Learning Scientist to innovate AI-based drug discovery with multimodal models.

Iambic Therapeutics Boston Office Published 2 weeks ago
Flexible on stack
Inductive Bio

Join Inductive Bio as a software engineer to build AI tools that accelerate drug discovery.

Inductive Bio New York City, San Francisco, or Boston Published 5 months ago
baseten

Join Baseten as an Infrastructure Software Engineer to build and maintain components of our ML inference platform for AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack