"inference serving systems" Jobs

285 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 285 results

baseten

Join Baseten as a Site Reliability Engineer to enhance the reliability of our multi-cloud Kubernetes infrastructure.

baseten San Francisco Published 4 months ago
Flexible on stack 70% coding
Harvey AI

Lead the design and development of systems powering AI requests at Harvey, collaborating with multiple teams to ensure reliability and efficiency.

Harvey AI San Francisco $236k–$290k/yr Published 1 month ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to design and optimize large-scale AI training and inference clusters.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Harvey

Lead the design and development of systems powering AI requests at Harvey, a fast-scaling company in the legal tech space.

Harvey San Francisco $231k–$340k/yr Published 1 week ago
Flexible on stack
baseten

Join Baseten as a Cloud Platform Engineer to build scalable infrastructure for deploying machine learning models.

baseten San Francisco Published 11 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Inferact

Join Inferact as an IT Support & Operations Engineer to enhance internal technology and security for a growing AI startup.

Inferact San Francisco $125k–$170k/yr Published 13 hours ago
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
Hilbert

Join Hilbert as an AI Engineer to build production-grade AI systems that drive enterprise outcomes in a fast-paced startup environment.

Hilbert San Francisco Published 19 hours ago
Flexible on stack 70% coding
baseten

Join Baseten as a Forward Deployed Engineer to solve complex AI challenges for leading companies.

baseten San Francisco Published 3 weeks ago
Flexible on stack
baseten

Join Baseten as a lead Software Engineer to own and develop production-grade Voice AI systems that impact daily lives.

baseten San Francisco Published 4 months ago
Flexible on stack
Inferact

Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack
Twelve Labs

Lead the development of next-generation multimodal models at Twelve Labs, impacting thousands of customers worldwide.

Twelve Labs Seoul, South Korea Published 1 week ago
Flexible on stack
Cartesia

Join Cartesia as a Software Engineer to design and build scalable AI model inference systems in a collaborative, in-office environment.

Cartesia *HQ - San Francisco, CA Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 1 week ago
Twelve Labs

Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 5 months ago
Flexible on stack
baseten

Join Baseten as a senior software engineer to develop cutting-edge AI training products and enhance user workflows.

baseten San Francisco Published 7 months ago
Flexible on stack