"inference optimization" Jobs

588 open tech roles matching “inference optimization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 588 results

Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Iambic Therapeutics

Join Iambic Therapeutics as a Machine Learning Scientist to innovate AI-based drug discovery with multimodal models.

Iambic Therapeutics UK Office Published 2 weeks ago
Flexible on stack
Imprint

Deliver analytical projects that influence product decisions and marketing campaigns in a fast-paced startup environment.

Imprint New York City Published 1 month ago
Flexible on stack
baseten

Join Baseten as a lead Software Engineer to own and develop production-grade Voice AI systems that impact daily lives.

baseten San Francisco Published 4 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems and engage with enterprise customers on generative AI solutions.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Twitch

Join Twitch's Monetization team as a Data Scientist to drive product decisions through rigorous analysis and causal inference.

Twitch San Francisco, CA $136k–$212.8k/yr Published 2 months ago
Flexible on stack
Twitch

Join Twitch's Monetization team as a Data Scientist to drive product decisions through causal analysis and experimentation.

Twitch Seattle, WA $136k–$212.8k/yr Published 2 months ago
Flexible on stack
Twitch

Join Twitch's Monetization team as a Data Scientist to apply causal inference methods and optimize revenue for creators.

Twitch New York City $136k–$212.8k/yr Published 2 months ago
Flexible on stack
Thumbtack

Join Thumbtack as an Applied Scientist to drive machine learning projects that enhance pro acquisition and marketplace efficiency.

Thumbtack Remote, Ontario CA$161.5k–CA$209k/yr Published 4 months ago
Flexible on stack
Omnifold

Lead a research team at Omnifold to develop advanced forecasting and optimization models in a startup environment.

Omnifold San Francisco, United States Published 1 day ago
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack