"inference serving" Jobs

905 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 905 results

Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft San Francisco, CA $148k–$185k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Inflection AI

Lead the development of Inflection's realtime Voice AI stack, shaping emotionally intelligent AI for enterprise voice interactions.

Inflection AI Palo Alto, California, United States $400k–$550k/yr Published 2 months ago
Inflection AI
Inflection AI Palo Alto, California, United States $350k–$500k/yr Published 4 months ago
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Instacart

Join Instacart as a Senior Marketing Decision Scientist II to drive data-driven growth and optimize marketing performance.

Instacart Canada - Remote (ON, AB, BC, or NS Only) CA$168k–CA$177.5k/yr Published 3 months ago
Flexible on stack
baseten

Join Baseten as a senior software engineer to develop cutting-edge AI training products and enhance user workflows.

baseten San Francisco Published 7 months ago
Flexible on stack
ChipAgents

Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.

ChipAgents San Jose $150k–$350k/yr Published 3 months ago
Flexible on stack
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Peregrine

Lead the development of AI-powered features for an end-to-end intelligence platform in public safety.

Peregrine San Francisco, CA $225k–$320k/yr Published 7 months ago
Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft New York, NY $148k–$185k/yr Published 3 months ago
Flexible on stack
Twelve Labs

Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 5 months ago
Flexible on stack
Twelve Labs

Lead the development of next-generation multimodal models at Twelve Labs, impacting thousands of customers worldwide.

Twelve Labs Seoul, South Korea Published 1 week ago
Flexible on stack
Patreon

Join Patreon as a Senior Machine Learning Engineer to architect and maintain high-throughput ML infrastructure for creator discovery.

Patreon New York Published 3 weeks ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA $315k–$560k/yr Published 10 months ago
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems and engage with enterprise customers on generative AI solutions.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 1 week ago