"inference serving" Jobs

391 open tech roles matching “inference serving”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 391 results

Hippocratic AI

Own the serving infrastructure for healthcare AI, optimizing LLM inference systems to enhance patient experiences.

Hippocratic AI Menlo Park, CA Published 3 weeks ago
Flexible on stack
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Roboflow

Join Roboflow as a Machine Learning Engineer to enhance our inference engine and contribute to impactful computer vision projects.

Roboflow NY, SF or Remote Published 2 months ago
Flexible on stack
Coreweave

Lead complex, cross-functional programs for inference platform delivery at a rapidly growing AI cloud company.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA $198k–$264k/yr Published 2 months ago
Together AI

Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.

Together AI Singapore Published 1 month ago
Flexible on stack 70% coding
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago
Fundamental

Join Fundamental as a Model Serving Engineer to optimize and scale the NEXUS model for enterprise decision-making.

Fundamental Europe Published 5 months ago
Flexible on stack
Chai Discovery

Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.

Chai Discovery San Francisco office Published 9 months ago
Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Instacart

Lead the design and development of core ML models for Instacart’s ads ecosystem in a fully remote role.

Instacart United States - Remote $201k–$253.5k/yr Published 3 months ago
Flexible on stack
Together AI

Join Together AI as a Technical Support Engineer to tackle complex technical challenges in a fast-paced AI environment.

Together AI Remote $160k–$230k/yr Published 1 month ago
Flexible on stack
Coreweave

Lead a team of engineers to build and operate CoreWeave's next-generation Kubernetes-native inference platform.

Coreweave Bellevue, WA - US $188k–$303k/yr Published 8 months ago
Heavy meetings
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $165k–$242k/yr Published 11 months ago