"inference serving" Jobs

350 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 350 results

Inferact

Join Inferact as a Product Marketing Manager to enhance vLLM's presence in the AI inference space through strategic marketing and community engagement.

Inferact San Francisco Published 1 month ago
Perplexity AI

Lead the economics of AI products at Perplexity AI, optimizing model spend and driving pricing and margin decisions.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Post-Training Research Scientist to advance AI research and collaborate on impactful projects.

baseten San Francisco Published 6 months ago
Faire

Own the measurement and optimization of long-term value in a hybrid role at Faire, a tech-driven wholesale platform.

Faire New York City, NY; San Francisco, CA $246.5k–$339k/yr Published 3 days ago
Flexible on stack
baseten

Join Baseten as a Forward Deployed Engineer to solve complex AI challenges for leading companies.

baseten San Francisco Published 3 weeks ago
Flexible on stack
Databricks

Join Databricks as an Applied AI Engineer to build personalized learning experiences using machine learning and knowledge representation.

Databricks United States $139k–$191.1k/yr Published 2 months ago
Flexible on stack
baseten

Join Baseten as a lead Software Engineer to own and develop production-grade Voice AI systems that impact daily lives.

baseten San Francisco Published 4 months ago
Flexible on stack
Mixpanel

Join Mixpanel as a Senior Data Scientist to drive AI-powered product insights and analytics for thousands of companies.

Mixpanel San Francisco, US (Hybrid) $216k–$254k/yr Published 1 month ago
Flexible on stack
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Inferact

Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack
Inferact

Join Inferact as an IT Support & Operations Engineer to enhance internal technology and security for a growing AI startup.

Inferact San Francisco $125k–$170k/yr Published 3 days ago
Strava

Lead data science efforts at Strava to connect product and marketing initiatives with measurable business outcomes.

Strava Strava SF Published 1 week ago
Flexible on stack
Cartesia

Join Cartesia as a Software Engineer to shape data infrastructure for cutting-edge AI models in a collaborative, in-office environment.

Cartesia *HQ - San Francisco, CA Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Cloud Platform Engineer to build scalable infrastructure for deploying machine learning models.

baseten San Francisco Published 11 months ago
Flexible on stack
baseten

Join Baseten as a Site Reliability Engineer to enhance the reliability of our multi-cloud Kubernetes infrastructure.

baseten San Francisco Published 4 months ago
Flexible on stack 70% coding
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft San Francisco, CA $148k–$185k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a Technical Staff member to enhance our AI inference engine with cutting-edge technologies.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings