"llm inference" Jobs

451 open tech roles matching “llm inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 451 results

Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
Anthropic

Join Anthropic's Inference team to build and maintain systems that serve AI models to millions of users worldwide.

Anthropic Ontario, CAN Published 3 weeks ago
Flexible on stack
Fundamental

Join Fundamental as an ML Researcher to tackle groundbreaking challenges in AI model development for enterprise decision-making.

Fundamental Barcelona Published 10 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
DeepL

Lead the Production Inference team at DeepL, focusing on performance-critical model serving systems in a fast-paced AI environment.

DeepL London Published 1 month ago
Heavy meetings
Anthropic
Anthropic San Francisco, CA $315k–$560k/yr Published 10 months ago
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.

Perplexity AI London Published 5 months ago
Flexible on stack
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models and AI applications.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 3 years ago
OpenRouter

Conduct original research on large language models to advance understanding and routing optimization at OpenRouter.

OpenRouter Remote (US) Published 2 months ago
Flexible on stack
Wispr Flow
ML Engineer Hybrid Visa

Join Wispr Flow as a ML Engineer to build a scalable voice interface for millions of users.

Wispr Flow San Francisco Published 1 year ago
Flexible on stack
Genesis Molecular AI

Join Genesis Molecular AI as an ML Research Engineer to develop cutting-edge foundation models for drug discovery.

Genesis Molecular AI San Mateo, CA Published 1 year ago
Flexible on stack 70% coding
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Anthropic

Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.

Anthropic San Francisco, CA | New York City, NY $350k–$850k/yr Published 3 days ago
Flexible on stack
Lilt

Join LILT as a Forward Deployed Engineer to integrate AI solutions for complex clients and enhance global communication.

Lilt London, UK Published 1 month ago
Flexible on stack