"llm inference systems" Jobs
392 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 392 results
Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models within a fully remote team in Switzerland.
Lead a multidisciplinary research team to advance large-scale machine learning efficiency at Databricks.
Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.
Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.
Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models and impact AI applications globally.
Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.
Join Anthropic's Inference team to design and maintain distributed systems that serve AI models to millions of users worldwide.
Join Anthropic's Inference team to design and maintain distributed systems serving AI models to millions globally.
Join Lyft as a Data Scientist to leverage causal inference for enhancing safety and customer care experiences.
Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.
Join Perplexity AI as an AI Infrastructure Engineer to build and optimize large-scale AI training and inference clusters.
Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.
Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.
Join Triomics as an MLOps & Data Engineer to build infrastructure for ML workflows in oncology, impacting patient outcomes.
Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI workloads.