"llm inference systems" Jobs

144 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 144 results

Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 4 days ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to design and optimize large-scale AI training and inference clusters.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago
Databricks

Lead a multidisciplinary research team to advance large-scale machine learning efficiency at Databricks.

Databricks Mountain View, California; San Francisco, California $270k–$340k/yr Published 3 months ago
Flexible on stack 60% coding
Anthropic
Anthropic San Francisco, CA $315k–$560k/yr Published 10 months ago
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Anthropic

Join Anthropic's Inference team to design and maintain distributed systems that serve AI models to millions of users worldwide.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 3 months ago
Flexible on stack
Anthropic

Join Anthropic's Inference team to design and maintain distributed systems serving AI models to millions globally.

Anthropic New York City, NY; San Francisco, CA | Seattle, WA $320k–$485k/yr Published 3 weeks ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
Anthropic

Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.

Anthropic San Francisco, CA | New York City, NY $350k–$850k/yr Published 5 days ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI workloads.

Databricks San Francisco, California $190k–$265k/yr Published 2 months ago
Flexible on stack
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
Anthropic

Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack