"inference serving systems" Jobs

283 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 283 results

baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Anthropic

Join Anthropic as a Staff Software Engineer to build scalable ML infrastructure for AI safety systems.

Anthropic San Francisco, CA $320k–$485k/yr Published 1 week ago
Flexible on stack
Together AI

Join Together AI as a Staff Software Engineer to build systems that automate GPU infrastructure management.

Together AI San Francisco $240k–$280k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Together AI

Design and deliver multi-petabyte storage systems for AI workloads at Together AI, optimizing performance and cost.

Together AI San Francisco $250k–$300k/yr Published 3 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer on the Observability team to enhance the reliability of AI product systems.

baseten San Francisco Published 1 month ago
Flexible on stack
Patreon

Join Patreon as a Senior Machine Learning Engineer to architect and maintain high-throughput ML infrastructure for creator discovery.

Patreon New York Published 1 month ago
Flexible on stack
Inductive Bio

Join Inductive Bio as a software engineer to build AI tools that accelerate drug discovery.

Inductive Bio New York City, San Francisco, or Boston Published 5 months ago
Sierra

Join Sierra as a Software Engineer, Infrastructure, to design and maintain core systems for our AI platform in a collaborative environment.

Sierra San Francisco, CA Published 3 months ago
Flexible on stack
Twelve Labs

Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
baseten

Join Baseten as a Post-Training Research Scientist to advance AI research and collaborate on impactful projects.

baseten San Francisco Published 6 months ago
Sesame

Join Sesame as a Backend Software Engineer to tackle complex challenges in building reliable, scalable systems for innovative voice agents.

Sesame San Francisco Published 3 weeks ago
Flexible on stack 70% coding
Mixpanel

Join Mixpanel as a Senior Data Scientist to drive AI-powered product insights and analytics for thousands of companies.

Mixpanel San Francisco, US (Hybrid) $216k–$254k/yr Published 1 month ago
Flexible on stack
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 4 months ago
Flexible on stack 60% coding
Perplexity AI

Join Perplexity AI as a staff software engineer to enhance our model serving platform for AI products and infrastructure.

Perplexity AI San Francisco Published 3 weeks ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI workloads.

Databricks San Francisco, California $190k–$265k/yr Published 2 months ago
Flexible on stack
Perplexity

Join Perplexity as a staff software engineer to enhance our model serving platform, impacting AI product development.

Perplexity San Francisco Published 3 weeks ago
Flexible on stack
Decagon

Design and operate data systems that power Decagon's AI products, ensuring high reliability and performance.

Decagon San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack
Faire

Own the measurement and optimization of long-term value in a hybrid role at Faire, a tech-driven wholesale platform.

Faire New York City, NY; San Francisco, CA $246.5k–$339k/yr Published 6 days ago
Flexible on stack