"inference optimization" Jobs

214 open tech roles matching “inference optimization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 214 results

baseten

Join Baseten as a lead Software Engineer to own and develop production-grade Voice AI systems that impact daily lives.

baseten San Francisco Published 4 months ago
Flexible on stack
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Twitch

Join Twitch's Monetization team as a Data Scientist to drive product decisions through rigorous analysis and causal inference.

Twitch San Francisco, CA $136k–$212.8k/yr Published 2 months ago
Flexible on stack
Twitch

Join Twitch's Monetization team as a Data Scientist to drive product decisions through causal analysis and experimentation.

Twitch Seattle, WA $136k–$212.8k/yr Published 2 months ago
Flexible on stack
Twitch

Join Twitch's Monetization team as a Data Scientist to apply causal inference methods and optimize revenue for creators.

Twitch New York City $136k–$212.8k/yr Published 2 months ago
Flexible on stack
Omnifold

Lead a research team at Omnifold to develop advanced forecasting and optimization models in a startup environment.

Omnifold San Francisco HQ Published 3 days ago
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
World Labs

Join World Labs as a Performance Engineer to optimize AI models for speed and efficiency in a cutting-edge research environment.

World Labs San Francisco $200k–$300k/yr Published 4 months ago
Flexible on stack 70% coding
Perplexity

Join Perplexity as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions in a fast-paced environment.

Perplexity San Francisco Published 1 week ago
baseten

Join Baseten as a Senior Frontend Engineer to craft user-friendly interfaces for AI systems in a collaborative environment.

baseten San Francisco Published 2 months ago
Perplexity AI

Join Perplexity AI as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions.

Perplexity AI San Francisco Published 1 week ago
Anthropic

Join Anthropic as a Staff Software Engineer to build scalable ML infrastructure for AI safety systems.

Anthropic San Francisco, CA $320k–$485k/yr Published 4 days ago
Flexible on stack
Lyft

Join Lyft as a Staff Applied Scientist to develop ML and optimization models that enhance pricing and ETA decisions.

Lyft San Francisco, CA $193.6k–$242k/yr Published 1 month ago
Flexible on stack