"ml model serving" Jobs

1271 open tech roles matching “ml model serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 1271 results

Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI UK £140k–£200k/yr Published 5 months ago
Flexible on stack
ChipAgents

Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.

ChipAgents San Jose $150k–$350k/yr Published 3 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 5 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Bayesian Health

Join Bayesian Health as a Staff Machine Learning Engineer to develop and deploy impactful ML models in healthcare.

Bayesian Health Remote Published 1 year ago
Flexible on stack 70% coding
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Sprinter Health

Join Sprinter Health as a Staff Machine Learning Engineer to build and lead the ML engineering function in a hybrid work environment.

Sprinter Health San Francisco, CA Published 1 month ago
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.

Inferact Remote Published 1 week ago
Flexible on stack
Airbnb

Lead the fine-tuning and optimization of LLMs to enhance AI products at Airbnb with a focus on customer support.

Airbnb United States $292k–$365k/yr Published 3 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic environment.

Inworld AI Serbia Published 5 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build CustomerLake, enhancing ML and AI personalization for businesses.

Databricks New York City, New York $192k–$260k/yr Published 2 months ago
Flexible on stack 60% coding
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Fundamental

Lead a team of MLOps engineers at an AI company transforming enterprise decision-making.

Fundamental Europe Published 2 months ago
Flexible on stack
HoneyBook

Join HoneyBook as a Senior Machine Learning Engineer to build and maintain AI-powered systems in a fast-paced environment.

HoneyBook Tel Aviv Published 3 months ago
Flexible on stack
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 3 days ago
Flexible on stack
Cantina

Join Cantina as an MLOps Engineer to build and scale inference infrastructure for generative audio models.

Cantina Remote (U.S. or Europe) $125k–$165k/yr Published 1 month ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve state-of-the-art voice models in a fully remote role.

Inworld AI Switzerland Published 5 months ago
Flexible on stack