"inference serving" Jobs

886 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 886 results

Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.

baseten San Francisco Published 5 months ago
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Product Manager to shape the future of AI inference and platform capabilities.

Fireworks AI San Mateo Published 3 weeks ago
AI-first team
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $139k–$204k/yr Published 7 months ago
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
ElevenLabs

Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.

ElevenLabs United Kingdom Published 2 weeks ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack
Applied Intuition

Design and implement software and machine learning components for behavior prediction and environmental interactions in a dynamic environment.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
Inferact
Head of Legal Hybrid Visa

Join Inferact as the first in-house legal hire to lead legal functions and support a fast-growing AI inference company.

Inferact San Francisco, CA, United States Published 3 days ago
Perplexity

Lead the financial strategy for AI products at Perplexity, optimizing model spend and driving pricing decisions.

Perplexity San Francisco Published 1 week ago
baseten

Join Baseten as an Infrastructure Software Engineer to build and maintain components of our ML inference platform for AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack
Inferact

Join Inferact as a Product Marketing Manager to enhance vLLM's presence in the AI inference space through strategic marketing and community engagement.

Inferact San Francisco Published 1 month ago
Volta

Lead product strategy for inference infrastructure and token-serving capabilities in a rapidly growing AI infrastructure company.

Volta Palo Alto, CA Published 1 month ago
Perplexity AI

Lead the economics of AI products at Perplexity AI, optimizing model spend and driving pricing and margin decisions.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Post-Training Research Scientist to advance AI research and collaborate on impactful projects.

baseten San Francisco Published 5 months ago
Inflection AI

Lead model training and post-training strategies for emotionally intelligent AI at Inflection AI.

Inflection AI Palo Alto, California, United States $400k–$550k/yr Published 2 months ago