"inference serving" Jobs

905 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 905 results

Applied Intuition

Join Applied Intuition as a Software Engineer to optimize application-layer software for embedded systems in autonomous driving.

Applied Intuition Sunnyvale Published 1 year ago
Motive

Join Motive as a Senior AI Platform Engineer to build and optimize production AI systems in a fully remote environment.

Motive India Published 1 day ago
Bluefish AI

Join Bluefish AI as a Senior/Staff Data Scientist to lead analytics and experimentation efforts in a fast-growing AI marketing platform.

Bluefish AI New York, Hybrid Published 3 months ago
Flexible on stack
Anthropic

Join Anthropic as a Demand Planning expert to optimize AI infrastructure capacity and ensure timely delivery across multiple platforms.

Anthropic San Francisco, CA | New York City, NY $320k–$405k/yr Published 1 month ago
Flexible on stack
Truecaller

Join Truecaller as a Senior Data Engineer to build and own the data foundation for recommendation and advertising ML systems.

Truecaller Stockholm, Sweden Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Harvey AI

Join Harvey AI as a Product Data Scientist to shape product success metrics and drive user insights in a fast-paced environment.

Harvey AI San Francisco $155k–$240k/yr Published 2 months ago
Flexible on stack AI-first team
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
World Labs

Join World Labs as a Performance Engineer to optimize AI models for speed and efficiency in a cutting-edge research environment.

World Labs San Francisco $200k–$300k/yr Published 4 months ago
Flexible on stack 70% coding
Perplexity

Join Perplexity as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions in a fast-paced environment.

Perplexity San Francisco Published 1 week ago
Perplexity AI

Join Perplexity AI as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions.

Perplexity AI San Francisco Published 1 week ago
Gusto

Lead the Sales Data Science & Analytics team to drive insights and forecasting for revenue growth at Gusto.

Gusto Denver, CO - Hybrid; New York, New York, United States; San Francisco, CA - Hybrid $218k–$255k/yr Published 11 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer on the Observability team to enhance the reliability of AI product systems.

baseten San Francisco Published 1 month ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models within a fully remote team in Switzerland.

Inworld AI Switzerland Published 7 months ago
Faire

Drive strategic decisions and analytics for fulfillment at Faire, a tech platform supporting local retailers.

Faire San Francisco, CA $197.5k–$271.5k/yr Published 2 months ago
AI-first team
Harvey AI

Lead a high-performing team to develop and manage the model infrastructure platform at Harvey AI.

Harvey AI San Francisco $272k–$355k/yr Published 1 month ago
Heavy meetings
Reddit

Build and optimize large-scale machine learning systems for recommendation and personalization at Reddit.

Reddit Remote - United States $190.8k–$267.1k/yr Published 1 month ago
Flexible on stack
Pinterest

Join Pinterest as a Data Scientist II to enhance ML capabilities and drive innovations in ML infrastructure.

Pinterest Palo Alto, CA, US; Remote, US $114.3k–$235.3k/yr Published 1 month ago
Flexible on stack