"inference serving" Jobs

352 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 352 results

Anthropic

Join Anthropic as a Staff Software Engineer to enhance deployment infrastructure for AI systems in a collaborative environment.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 2 months ago
Flexible on stack
Chai Discovery

Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.

Chai Discovery San Francisco office Published 9 months ago
Inferact

Join Inferact as a Site Reliability Engineer to enhance the reliability and performance of AI inference systems at scale.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack AI-first team
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
baseten

Join Baseten as a Data Engineer to build and scale the internal data platform for AI-driven decision-making.

baseten San Francisco Published 5 months ago
baseten

Join Baseten as an AI Inference Engineer to architect and deploy high-scale production AI applications while collaborating with customers.

baseten San Francisco Published 1 month ago
Flexible on stack 70% coding
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.

baseten San Francisco Published 5 months ago
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack
Inferact
Head of Legal Hybrid Visa

Join Inferact as the first in-house legal hire to lead legal functions and support a fast-growing AI inference company.

Inferact San Francisco Published 4 days ago
Perplexity

Lead the financial strategy for AI products at Perplexity, optimizing model spend and driving pricing decisions.

Perplexity San Francisco Published 1 week ago
baseten

Join Baseten as an Infrastructure Software Engineer to build and maintain components of our ML inference platform for AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack