"inference serving systems" Jobs
282 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 282 results
Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.
Join Baseten as an Infrastructure Software Engineer to build and maintain components of our ML inference platform for AI applications.
Join Baseten as an AI Inference Engineer to architect and deploy high-scale production AI applications while collaborating with customers.
Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.
Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.
Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.
Join Anthropic as a Staff Software Engineer to enhance deployment infrastructure for AI systems in a collaborative environment.
Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.
Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.
Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.
Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.
Join Databricks as an Applied AI Engineer to build personalized learning experiences using machine learning and knowledge representation.
Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.
Join Cartesia as a Software Engineer to shape data infrastructure for cutting-edge AI models in a collaborative, in-office environment.