"llm inference systems" Jobs

392 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 392 results

Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Iambic Therapeutics

Join Iambic Therapeutics as a Machine Learning Scientist to innovate AI-based drug discovery with multimodal models.

Iambic Therapeutics UK Office Published 2 weeks ago
Flexible on stack
Applied Intuition

Join Applied Intuition as an Embedded AI Engineer to develop on-device intelligence for Android Automotive platforms.

Applied Intuition Sunnyvale Published 5 months ago
Flexible on stack
Ambient

Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced machine learning models.

Ambient Redwood City Published 2 days ago
Flexible on stack 70% coding
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models and AI applications.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 3 years ago
Fireworks AI

Join Fireworks AI as a Software Engineer to design and build scalable infrastructure for generative AI systems.

Fireworks AI San Mateo Published 10 months ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 2 days ago
Flexible on stack
Iambic Therapeutics

Join Iambic Therapeutics as a Machine Learning Scientist to innovate AI-based drug discovery with multimodal models.

Iambic Therapeutics Boston Office Published 2 weeks ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.

Perplexity AI London Published 5 months ago
Flexible on stack
Ambient

Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced AI models.

Ambient Redwood City Published 2 months ago
Flexible on stack 70% coding
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to design and optimize large-scale AI training and inference clusters.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago
DeepL

Lead the Production Inference team at DeepL, focusing on performance-critical model serving systems in a fast-paced AI environment.

DeepL London Published 1 month ago
Heavy meetings