"inference optimization" Jobs
588 open tech roles matching “inference optimization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 588 results
Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.
Design and implement software and machine learning components for behavior prediction and environmental interactions in a dynamic environment.
Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.
Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.
Lead the economics of AI products at Perplexity AI, optimizing model spend and driving pricing and margin decisions.
Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.
Lead the design and development of core ML models for Instacart’s ads ecosystem in a fully remote role.
Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic environment.
Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.
Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.
Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.
Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.
Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.
Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve state-of-the-art voice models in a fully remote role.
Join Perplexity AI as an AI Inference Engineer to optimize and develop our inference engine for various model architectures.
Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.
Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working directly with hardware vendors.
Join Inworld AI as a Lead Machine Learning Engineer to optimize and serve state-of-the-art voice models in a dynamic startup environment.
Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.