"llm inference systems" Jobs

392 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 392 results

Together AI

Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.

Together AI Singapore Published 1 month ago
Flexible on stack 70% coding
Fireworks AI

Design and optimize infrastructure for large-scale AI model training at a leading generative AI company.

Fireworks AI San Mateo Published 1 month ago
Flexible on stack
MaintainX

Lead the technical direction for predictive maintenance and asset intelligence initiatives at MaintainX, leveraging deep ML expertise.

MaintainX San Francisco Published 1 month ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Machine Learning Engineer to enhance search quality through innovative ranking solutions.

Perplexity AI Belgrade Published 1 month ago
Cylake

Join a small team to build state-of-the-art AI capabilities for Cylake's next-generation cybersecurity platform.

Cylake Sunnyvale $150k–$250k/yr Published 1 month ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to enhance deployment infrastructure for AI systems in a collaborative environment.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Adaptive Security

Join Adaptive as a Founding ML Engineer to build and define ML capabilities for AI-powered cybersecurity solutions.

Adaptive Security NYC Published 1 week ago
Flexible on stack 70% coding
Perplexity AI

Join Perplexity AI as a Senior Applied AI Engineer to shape agent capabilities and enhance user experiences with cutting-edge AI technologies.

Perplexity AI San Francisco Published 3 days ago
Cloudflare

Join Cloudflare as a Senior Systems Engineer to build core AI Gateway systems for high-volume inference traffic.

Cloudflare In-Office Published 1 month ago
Perplexity

Join Perplexity as a staff Applied AI Engineer to shape agent capabilities and enhance user experiences with cutting-edge AI technologies.

Perplexity San Francisco Published 3 days ago
Illumio

Architect high-scale distributed systems and lead the development of autonomous AI agents in a dynamic cybersecurity environment.

Illumio HQ - Sunnyvale (Office) Published 3 months ago
Flexible on stack
Freenome

Join Freenome as a Senior Machine Learning Engineer to develop and optimize deep learning pipelines for cancer detection.

Freenome Remote $173.8k–$246.8k/yr Published 1 month ago
Flexible on stack
Omnifold

Lead a research team at Omnifold to develop advanced forecasting and optimization models in a startup environment.

Omnifold San Francisco HQ Published 2 days ago
Harvey AI

Lead the design and development of systems powering AI requests at Harvey, collaborating with multiple teams to ensure reliability and efficiency.

Harvey AI San Francisco $236k–$290k/yr Published 1 month ago
Flexible on stack
Twelve Labs

Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.

Twelve Labs Seoul, South Korea Published 3 weeks ago
Flexible on stack
Cloudflare

Join Cloudflare as a Senior Machine Learning Engineer to optimize and productionize ML models for a global serverless inference platform.

Cloudflare Hybrid Published 2 months ago
Flexible on stack
Twelve Labs

Drive research on Pegasus's complex problems in a hybrid role at a growing AI company focused on video understanding.

Twelve Labs Seoul, South Korea Published 5 months ago
Flexible on stack
PagerDuty

Join PagerDuty as a junior AI/ML Engineer to build and ship AI systems at scale, collaborating with senior engineers.

PagerDuty Lisbon Published 3 weeks ago
Flexible on stack