"inference serving" Jobs
350 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 350 results
Join Inferact as a Product Marketing Manager to enhance vLLM's presence in the AI inference space through strategic marketing and community engagement.
Lead the economics of AI products at Perplexity AI, optimizing model spend and driving pricing and margin decisions.
Join Baseten as a Post-Training Research Scientist to advance AI research and collaborate on impactful projects.
Own the measurement and optimization of long-term value in a hybrid role at Faire, a tech-driven wholesale platform.
Join Baseten as a Forward Deployed Engineer to solve complex AI challenges for leading companies.
Join Databricks as an Applied AI Engineer to build personalized learning experiences using machine learning and knowledge representation.
Join Mixpanel as a Senior Data Scientist to drive AI-powered product insights and analytics for thousands of companies.
Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.
Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.
Join Inferact as an IT Support & Operations Engineer to enhance internal technology and security for a growing AI startup.
Lead data science efforts at Strava to connect product and marketing initiatives with measurable business outcomes.
Join Cartesia as a Software Engineer to shape data infrastructure for cutting-edge AI models in a collaborative, in-office environment.
Join Baseten as a Cloud Platform Engineer to build scalable infrastructure for deploying machine learning models.
Join Baseten as a Site Reliability Engineer to enhance the reliability of our multi-cloud Kubernetes infrastructure.
Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.
Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.
Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.
Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.