"llm inference systems" Jobs

19 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AWS. Every listing is re-checked daily and closed roles are removed.

Showing 19 of 19 results

Perplexity AI

Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.

Perplexity AI London Published 5 months ago
Flexible on stack
DeepL

Lead the Production Inference team at DeepL, focusing on performance-critical model serving systems in a fast-paced AI environment.

DeepL London Published 1 month ago
Heavy meetings
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to build and optimize large-scale AI training and inference clusters.

Perplexity AI London Published 5 months ago
Flexible on stack
Lilt

Lead complex benchmark transformations and data deliverables in a senior role focused on applied AI at Lilt.

Lilt London, UK Published 1 week ago
Flexible on stack
Lilt

Join LILT as a Forward Deployed Engineer to integrate AI solutions for complex clients and enhance global communication.

Lilt London, UK Published 1 month ago
Flexible on stack
DeepL

Lead research on fine-tuning and steerability of LLM-based translation models in a collaborative AI-focused environment.

DeepL London Published 1 month ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Member of Technical Staff to build innovative AI solutions on a large inference platform.

Fireworks AI London Published 6 days ago
Flexible on stack
DeepL

Join DeepL as a Senior Software Engineer to build innovative real-time voice translation solutions in a dynamic, cross-functional team.

DeepL London Published 1 month ago
Flexible on stack
Fireworks AI

Join Fireworks AI as an AI Field Engineer to build production systems for generative AI with large organizations across EMEA.

Fireworks AI London Published 1 month ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as an Applied Machine Learning Engineer to bridge AI research and real-world applications in a collaborative environment.

Fireworks AI London Published 1 week ago
Flexible on stack
DeepL

Lead scientific innovation in speech and translation models for real-time voice products at DeepL.

DeepL London Published 1 month ago
Flexible on stack 70% coding
Perplexity AI

Lead the establishment and growth of Perplexity's London office, shaping its engineering culture and technical direction.

Perplexity AI London Published 1 month ago
Perplexity

Lead the establishment and growth of Perplexity's London office, shaping its engineering culture and technical direction.

Perplexity London Published 1 month ago
Together AI

Join Together AI to build production AI agents and foundational systems for one of the largest GPU fleets in the world.

Together AI London Published 1 week ago
Flexible on stack
Magentic

Join Magentic as a Back-end Engineer to build scalable AI-driven backend services for enterprise supply chains.

Magentic London £125k–£140k/yr Published 3 months ago
Flexible on stack