"inference serving" Jobs

886 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 886 results

Coreweave

Lead complex, cross-functional programs for inference platform delivery at a rapidly growing AI cloud company.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA $198k–$264k/yr Published 2 months ago
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI UK £140k–£200k/yr Published 5 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic environment.

Inworld AI Serbia Published 5 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
SpaceX

Join SpaceX as a Software Engineer to develop high-performance AI inference systems for mission-critical applications.

SpaceX Palo Alto, CA $135k–$175k/yr Published 3 weeks ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Anthropic

Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.

Anthropic San Francisco, CA | New York City, NY $350k–$850k/yr Published 2 days ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Pika

Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.

Pika Palo Alto HQ Published 2 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 5 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Together AI

Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.

Together AI Singapore Published 1 month ago
Flexible on stack 70% coding
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve state-of-the-art voice models in a fully remote role.

Inworld AI Switzerland Published 5 months ago
Flexible on stack
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago