"llm inference systems" Jobs
392 open tech roles matching “llm inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 392 results
Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.
Join Iambic Therapeutics as a Machine Learning Scientist to innovate AI-based drug discovery with multimodal models.
Join Applied Intuition as an Embedded AI Engineer to develop on-device intelligence for Android Automotive platforms.
Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced machine learning models.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models and AI applications.
Join Fireworks AI as a Software Engineer to design and build scalable infrastructure for generative AI systems.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join Iambic Therapeutics as a Machine Learning Scientist to innovate AI-based drug discovery with multimodal models.
Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.
Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.
Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.
Design and optimize AI infrastructure for real-time intelligence at Ambient.ai, enhancing security through advanced AI models.
Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.
Join Perplexity AI as an AI Infrastructure Engineer to design and optimize large-scale AI training and inference clusters.
Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.
Lead the Production Inference team at DeepL, focusing on performance-critical model serving systems in a fast-paced AI environment.