"inference serving systems" Jobs

735 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 735 results

Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
Applied Intuition

Lead the perception model team for autonomous vehicles at a rapidly growing AI infrastructure company.

Applied Intuition Sunnyvale Published 3 months ago
Flexible on stack
Twelve Labs

Lead the development of next-generation multimodal models at Twelve Labs, impacting thousands of customers worldwide.

Twelve Labs Seoul, South Korea Published 1 week ago
Flexible on stack
Cartesia

Join Cartesia as a Software Engineer to design and build scalable AI model inference systems in a collaborative, in-office environment.

Cartesia *HQ - San Francisco, CA Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Inferact

Join Inferact as a cloud orchestration engineer to build reliable systems for AI model deployment at scale.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 1 week ago
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to build and optimize large-scale AI training and inference clusters.

Perplexity AI London Published 5 months ago
Flexible on stack
Applied Intuition

Join Applied Intuition as an Embedded AI Engineer to develop on-device intelligence for Android Automotive platforms.

Applied Intuition Sunnyvale Published 5 months ago
Flexible on stack
Twelve Labs

Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 5 months ago
Flexible on stack
Volta

Lead product strategy for inference infrastructure and token-serving capabilities in a rapidly growing AI infrastructure company.

Volta Palo Alto, CA Published 1 month ago
Periodic Labs

Join Periodic Labs as an ML Systems Engineer to build and optimize large-scale training and reinforcement learning infrastructure.

Periodic Labs Menlo Park, CA $250k–$350k/yr Published 4 months ago
Flexible on stack
baseten

Join Baseten as a senior software engineer to develop cutting-edge AI training products and enhance user workflows.

baseten San Francisco Published 7 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Enterpret

Own the Graph Platform at Enterpret, leading data systems and architecture decisions for a customer feedback intelligence platform.

Enterpret Bengaluru, Onsite Published 3 weeks ago
Flexible on stack
Applied Intuition

Join Applied Intuition as a Senior Software Engineer to design and implement ML infrastructure for deep learning model training.

Applied Intuition Sunnyvale $215k–$285k/yr Published 3 years ago
Flexible on stack