"inference serving" Jobs
362 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 362 results
Join Anthropic's Inference team to build and maintain systems that serve AI models to millions of users worldwide.
Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.
Lead a team of engineers to optimize Anthropic's inference infrastructure for AI systems.
Own the inference systems that power frontier AI models in production and research at a tech-first startup.
Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.
Join Anthropic's Inference team to design and maintain distributed systems that serve AI models to millions of users worldwide.
Join Anthropic's Inference team to design and maintain distributed systems serving AI models to millions globally.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.
Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.
Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.
Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.
Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.
Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.
Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.
Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.