"inference serving systems" Jobs

735 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 735 results

Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $165k–$242k/yr Published 11 months ago
Together AI

Join Together AI as a Technical Support Engineer to tackle complex technical challenges in a fast-paced AI environment.

Together AI Remote $160k–$230k/yr Published 1 month ago
Flexible on stack
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago
Together AI

Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.

Together AI Singapore Published 1 month ago
Flexible on stack 70% coding
Cloudflare

Join Cloudflare as a Senior Systems Engineer to build core AI Gateway systems for high-volume inference traffic.

Cloudflare In-Office Published 1 month ago
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $139k–$204k/yr Published 7 months ago
Applied Intuition

Join Applied Intuition as an AI Performance Engineer to optimize large-scale machine learning workloads in a collaborative environment.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
ElevenLabs

Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.

ElevenLabs United Kingdom Published 2 weeks ago
Flexible on stack
baseten

Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.

baseten San Francisco Published 5 months ago
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
Sarvam AI

Own the architecture of Sarvam's vision models serving harness, ensuring high-quality document intelligence at national scale.

Sarvam AI Bengaluru Published 3 weeks ago
Flexible on stack 70% coding
Perplexity AI

Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.

Perplexity AI London Published 5 months ago
Flexible on stack
Coreweave

Lead a team of engineers to build and operate CoreWeave's next-generation Kubernetes-native inference platform.

Coreweave Bellevue, WA - US $188k–$303k/yr Published 8 months ago
Heavy meetings
Applied Intuition

Join Applied Intuition as a Perception Software Engineer to develop safety-critical perception systems for L4 autonomous trucks.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
Databricks

Join Databricks as an Applied AI Engineer to build personalized learning experiences using machine learning and knowledge representation.

Databricks United States $139k–$191.1k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack