"llm inference systems" Jobs

392 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 392 results

Inworld AI

Join Inworld AI as a Staff/Principal Software Engineer to develop cutting-edge backend systems for real-time voice models.

Inworld AI Vancouver, British Columbia, Canada CA$180k–CA$260k/yr Published 1 year ago
Flexible on stack 70% coding
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Twelve Labs

Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 1 week ago
krea.ai

Join Krea to build innovative AI tools in a hands-on role focused on supercomputing and distributed systems.

krea.ai San Francisco Published 5 months ago
Flexible on stack
OpenRouter

Join OpenRouter as a Software Engineer to scale systems handling millions of LLM requests daily in a fast-growing AI infrastructure company.

OpenRouter Remote (US) Published 4 months ago
Flexible on stack 70% coding
Coreweave

Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.

Coreweave Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA $188k–$275k/yr Published 6 months ago
Flexible on stack
Typeface

Lead the design and strategy for large-scale ML systems and generative AI at Typeface, influencing company-level initiatives.

Typeface Palo Alto, CA $230k–$260k/yr Published 4 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Member of Technical Staff to build innovative AI solutions on a large inference platform.

Fireworks AI London Published 6 days ago
Flexible on stack
DeepL

Join DeepL as a Senior Software Engineer to build innovative real-time voice translation solutions in a dynamic, cross-functional team.

DeepL London Published 1 month ago
Flexible on stack
baseten

Join the Base Labs Fellowship to conduct cutting-edge AI research with mentorship and funding in San Francisco.

baseten San Francisco $15k–$15k/mo Published 2 months ago
PagerDuty

Join PagerDuty as a Senior AI/ML Engineer to design and build AI-powered features for high-volume, real-time event streams.

PagerDuty Lisbon Published 3 weeks ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Member of Technical Staff to advance generative AI through foundational research and collaboration with top experts.

Fireworks AI San Mateo Published 1 month ago
Flexible on stack
Anthropic

Design and operate backend systems for Claude's safety systems, ensuring low latency and high reliability.

Anthropic San Francisco, CA $320k–$485k/yr Published 4 days ago
LangChain

Join LangChain as a Research Engineer to enhance the capabilities of the LangSmith Engine for AI agents.

LangChain New York, NY Published 1 month ago
Flexible on stack
Airbnb

Architect and maintain Airbnb’s end-to-end traffic classification ML systems to enhance bot detection and traffic integrity.

Airbnb United States $212k–$265k/yr Published 1 month ago
Hebbia AI

Join Hebbia AI as a Backend Engineer to build scalable solutions for advanced AI-driven investment analysis.

Hebbia AI NYC $160k–$350k/yr Published 1 year ago
Flexible on stack
OpenRouter

Join OpenRouter as a Data Scientist to enhance AI interactions for millions of developers and optimize product experiences.

OpenRouter Remote (US) Published 1 month ago
Flexible on stack
Sila Nanotechnologies

Join Sila as a Staff Applied ML Engineer to build intelligence systems for manufacturing operations and drive impactful engineering solutions.

Sila Nanotechnologies Alameda, CA $151k–$177.5k/yr Published 5 months ago
Flexible on stack