"inference serving systems" Jobs

33 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, Rust. Every listing is re-checked daily and closed roles are removed.

Showing 13 of 33 results

Fireworks AI

Join Fireworks AI as an Applied Machine Learning Engineer to bridge AI research and real-world applications in a collaborative environment.

Fireworks AI London Published 1 week ago
Flexible on stack
Perplexity AI

Lead the establishment and growth of Perplexity's London office, shaping its engineering culture and technical direction.

Perplexity AI London Published 1 month ago
Perplexity

Lead the establishment and growth of Perplexity's London office, shaping its engineering culture and technical direction.

Perplexity London Published 1 month ago
Neko Health

Join Neko Health as an ML Ops Engineer to lead the ML infrastructure for preventive care and early detection.

Neko Health London Published 2 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Senior Specialist Field Engineer to lead innovation in AI cloud infrastructure and networking technologies.

Coreweave London, UK £98k–£130k/yr Published 3 months ago
Flexible on stack
DeepL

Join DeepL as a Senior Software Engineer to build innovative real-time voice translation solutions in a dynamic, cross-functional team.

DeepL London Published 1 month ago
Flexible on stack
Volta

Lead a team of senior network platform engineers to design and operate large scale GPU compute infrastructure for AI workloads.

Volta Palo Alto, CA Published 1 month ago
Flexible on stack
Fireworks AI

Drive sourced pipeline and revenue through the Microsoft Azure channel as a senior sales operator at Fireworks AI.

Fireworks AI London Published 1 month ago
Volta

Join Volta as a Platform Engineer to build and operate large-scale GPU compute infrastructure for AI workloads.

Volta Palo Alto, CA Published 1 month ago
Flexible on stack
Volta

Join Volta as a Platform Engineer to build and operate large-scale GPU compute infrastructure for AI workloads.

Volta London, UK Published 1 month ago
Flexible on stack
DeepL

Lead scientific innovation in speech and translation models for real-time voice products at DeepL.

DeepL London Published 1 month ago
Flexible on stack 70% coding
Volta

Lead a platform engineering team at Volta, focusing on building and operating large-scale GPU compute infrastructure for AI workloads.

Volta Palo Alto, CA Published 1 month ago
Flexible on stack
Volta

Lead a platform engineering team at Volta, focusing on building and operating large-scale GPU compute infrastructure for AI workloads.

Volta London, UK Published 1 month ago
Flexible on stack