"inference serving" Jobs

353 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 353 results

Twelve Labs

Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.

Twelve Labs Seoul, South Korea Published 3 weeks ago
Flexible on stack
Mercury

Build and operate real-time inference services for risk decisioning in a fast-growing fintech startup.

Mercury San Francisco, CA, New York, NY, Portland, OR, or Remote within Canada or United States $166.6k–$208.3k/yr Published 5 days ago
Flexible on stack
Faire

Join Faire as a Senior Applied AI/ML Scientist to drive brand growth through innovative machine learning solutions.

Faire San Francisco, CA $211k–$290.5k/yr Published 3 days ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to build an AI developer platform that enhances productivity for engineers.

baseten San Francisco Published 1 month ago
Flexible on stack
Cartesia

Join Cartesia as a Software Engineer to design and build scalable AI model inference systems in a collaborative, in-office environment.

Cartesia *HQ - San Francisco, CA Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to build scalable ML infrastructure for AI safety systems.

Anthropic San Francisco, CA $320k–$485k/yr Published 5 days ago
Flexible on stack
baseten

Join Baseten as a Sr. Analyst in Revenue Strategy & Operations to shape GTM strategies for AI infrastructure.

baseten San Francisco Published 1 week ago
baseten

Join Baseten as a Forward Deployed Engineer to architect and deploy high-scale AI applications while collaborating with customers.

baseten San Francisco Published 2 years ago
Flexible on stack 70% coding
Harvey

Lead the design and development of systems powering AI requests at Harvey, a fast-scaling company in the legal tech space.

Harvey San Francisco $231k–$340k/yr Published 1 week ago
Flexible on stack
Inferact

Lead HR and People Operations at Inferact, scaling infrastructure in a fast-paced startup environment.

Inferact San Francisco $180k–$250k/yr Published 5 days ago
Faire

Join Faire as a Senior Applied AI/ML Scientist to drive retailer growth through innovative machine learning solutions.

Faire Kitchener-Waterloo, ON; Toronto, ON CA$180k–CA$247.5k/yr Published 1 month ago
Flexible on stack
Anthropic

Join Anthropic as a Demand Planning expert to optimize AI infrastructure capacity and ensure timely delivery across multiple platforms.

Anthropic San Francisco, CA | New York City, NY $320k–$405k/yr Published 1 month ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Harvey AI

Join Harvey AI as a Product Data Scientist to shape product success metrics and drive user insights in a fast-paced environment.

Harvey AI San Francisco $155k–$240k/yr Published 2 months ago
Flexible on stack AI-first team
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Perplexity

Join Perplexity as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions in a fast-paced environment.

Perplexity San Francisco Published 1 week ago
Perplexity AI

Join Perplexity AI as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions.

Perplexity AI San Francisco Published 1 week ago
Gusto

Lead the Sales Data Science & Analytics team to drive insights and forecasting for revenue growth at Gusto.

Gusto Denver, CO - Hybrid; New York, New York, United States; San Francisco, CA - Hybrid $218k–$255k/yr Published 11 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer on the Observability team to enhance the reliability of AI product systems.

baseten San Francisco Published 1 month ago
Flexible on stack