"inference infrastructure" Jobs

88 open tech roles matching “inference infrastructure”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 88 results

Anthropic

Lead a team of engineers to optimize Anthropic's inference infrastructure for AI systems.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$625k/yr Published 1 week ago
Flexible on stack Heavy meetings
Volta

Lead product strategy for inference infrastructure and token-serving capabilities in a rapidly growing AI infrastructure company.

Volta Palo Alto, CA Published 1 month ago
Omnifold

Lead the infrastructure team at Omnifold, focusing on AI model training and deployment in a fast-paced startup environment.

Omnifold San Francisco HQ Published 6 months ago
Flexible on stack
Anthropic

Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.

Anthropic San Francisco, CA $500k–$850k/yr Published 1 month ago
Flexible on stack
baseten

Join Baseten as a lead Software Engineer to own and develop production-grade Voice AI systems that impact daily lives.

baseten San Francisco Published 4 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic environment.

Inworld AI Serbia Published 5 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic startup environment.

Inworld AI Germany Published 5 months ago
Flexible on stack
Abridge

Lead product strategy for foundational models and post-training at a growing healthcare AI startup in San Francisco.

Abridge SF Office Published 2 weeks ago
Peregrine

Lead the development of AI-powered features for an end-to-end intelligence platform in public safety.

Peregrine San Francisco, CA $225k–$320k/yr Published 7 months ago
baseten

Lead the Runtime Fabric team at Baseten to build container runtimes tailored for AI inference workloads.

baseten San Francisco Published 3 months ago
Flexible on stack Heavy meetings
Inferact

Lead HR and People Operations at Inferact, scaling infrastructure in a fast-paced startup environment.

Inferact San Francisco $180k–$250k/yr Published 3 days ago
DeepL

Lead the Production Inference team at DeepL, focusing on performance-critical model serving systems in a fast-paced AI environment.

DeepL London Published 1 month ago
Heavy meetings
Perplexity AI

Join Perplexity AI as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions.

Perplexity AI San Francisco Published 1 week ago
Perplexity

Join Perplexity as a Strategic Finance Lead to optimize GPU compute investments and drive capacity decisions in a fast-paced environment.

Perplexity San Francisco Published 1 week ago
Applied Intuition

Lead the perception model team for autonomous vehicles at a rapidly growing AI infrastructure company.

Applied Intuition Sunnyvale Published 3 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Lead Research Scientist to innovate in real-time voice models and AI applications.

Inworld AI Serbia Published 5 months ago
baseten

Lead a team of cloud platform engineers to build scalable and reliable infrastructure for AI products at Baseten.

baseten San Francisco Published 3 months ago
Flexible on stack Heavy meetings
Inworld AI

Join Inworld AI as a Lead Research Scientist to innovate in real-time voice models and impact AI applications globally.

Inworld AI Germany Published 5 months ago
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Reflection AI

Lead the post-training and evaluation capabilities for large language models in a dynamic AI research lab.

Reflection AI New York, NY Published 10 months ago