"llm inference systems" Jobs
144 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 144 results
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.
Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.
Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.
Join Perplexity AI as an AI Infrastructure Engineer to design and optimize large-scale AI training and inference clusters.
Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.
Lead a multidisciplinary research team to advance large-scale machine learning efficiency at Databricks.
Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.
Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.
Join Anthropic's Inference team to design and maintain distributed systems that serve AI models to millions of users worldwide.
Join Anthropic's Inference team to design and maintain distributed systems serving AI models to millions globally.
Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.
Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.
Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.
Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI workloads.
Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.
Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.