"inference serving" Jobs

886 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 886 results

Fundamental

Join Fundamental as a Model Serving Engineer to optimize and scale the NEXUS model for enterprise decision-making.

Fundamental Europe Published 5 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to enhance deployment infrastructure for AI systems in a collaborative environment.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 2 months ago
Flexible on stack
Chai Discovery

Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.

Chai Discovery San Francisco office Published 9 months ago
Inferact

Join Inferact as a Site Reliability Engineer to enhance the reliability and performance of AI inference systems at scale.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack AI-first team
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Inworld AI

Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic startup environment.

Inworld AI Germany Published 5 months ago
Flexible on stack
Instacart

Lead the design and development of core ML models for Instacart’s ads ecosystem in a fully remote role.

Instacart United States - Remote $201k–$253.5k/yr Published 3 months ago
Flexible on stack
baseten

Join Baseten as a Data Engineer to build and scale the internal data platform for AI-driven decision-making.

baseten San Francisco Published 5 months ago
Together AI

Join Together AI as a Technical Support Engineer to tackle complex technical challenges in a fast-paced AI environment.

Together AI Remote $160k–$230k/yr Published 1 month ago
Flexible on stack
Coreweave

Lead a team of engineers to build and operate CoreWeave's next-generation Kubernetes-native inference platform.

Coreweave Bellevue, WA - US $188k–$303k/yr Published 8 months ago
Heavy meetings
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as an AI Inference Engineer to architect and deploy high-scale production AI applications while collaborating with customers.

baseten San Francisco Published 1 month ago
Flexible on stack 70% coding
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $92k–$135k/yr Published 10 months ago
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $165k–$242k/yr Published 11 months ago