"llm inference" Jobs

165 open tech roles matching “llm inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 165 results

Reflection AI

Lead large-scale AI infrastructure engagements with governments and enterprises, shaping complex partnerships and driving strategic outcomes.

Reflection AI San Francisco, CA Published 2 weeks ago
baseten

Join the Base Labs Fellowship to conduct cutting-edge AI research with mentorship and funding in San Francisco.

baseten San Francisco $15k–$15k/mo Published 2 months ago
PagerDuty

Join PagerDuty as a junior AI/ML Engineer to build and ship AI systems at scale, collaborating with senior engineers.

PagerDuty Lisbon Published 3 weeks ago
Flexible on stack
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Perplexity

Join Perplexity as a staff Applied AI Engineer to shape agent capabilities and enhance user experiences with cutting-edge AI technologies.

Perplexity San Francisco Published 4 days ago
Perplexity AI

Join Perplexity AI as a Senior Applied AI Engineer to shape agent capabilities and enhance user experiences with cutting-edge AI technologies.

Perplexity AI San Francisco Published 4 days ago
Twelve Labs

Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.

Twelve Labs Seoul, South Korea Published 3 weeks ago
Flexible on stack
HappyRobot

Join HappyRobot as a Machine Learning Engineer to build AI models for human-like conversations and shape the future of AI infrastructure.

HappyRobot San Francisco Published 1 month ago
Flexible on stack
Abridge

Lead product strategy for foundational models and post-training at a growing healthcare AI startup in San Francisco.

Abridge SF Office Published 2 weeks ago
Twelve Labs

Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 5 months ago
Flexible on stack
Anthropic

Design and operate backend systems for Claude's safety systems, ensuring low latency and high reliability.

Anthropic San Francisco, CA $320k–$485k/yr Published 5 days ago
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 1 week ago
PagerDuty

Join PagerDuty as a Senior AI/ML Engineer to design and build AI-powered features for high-volume, real-time event streams.

PagerDuty Lisbon Published 3 weeks ago
Flexible on stack
baseten

Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.

baseten San Francisco Published 5 months ago
krea.ai

Join Krea as an ML Researcher to train diffusion models for image and video generation in a creative AI-focused environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Sprinter Health

Join Sprinter Health as a Staff Machine Learning Engineer to build and lead the ML engineering function in a hybrid work environment.

Sprinter Health San Francisco, CA Published 1 month ago
baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Harvey

Lead the design and development of systems powering AI requests at Harvey, a fast-scaling company in the legal tech space.

Harvey San Francisco $231k–$340k/yr Published 1 week ago
Flexible on stack
LangChain

Join LangChain as a Research Engineer to enhance the capabilities of the LangSmith Engine for AI agents.

LangChain New York, NY Published 1 month ago
Flexible on stack