"inference serving systems" Jobs

735 open tech roles matching “inference serving systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 735 results

Sesame

Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.

Sesame San Francisco Published 1 year ago
Flexible on stack
Lyft

Join Lyft as a Data Scientist to leverage causal inference for enhancing safety and customer care experiences.

Lyft Toronto, Canada CA$108k–CA$135k/yr Published 1 month ago
Flexible on stack
Coreweave

Lead complex, cross-functional programs for inference platform delivery at a rapidly growing AI cloud company.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA $198k–$264k/yr Published 2 months ago
Applied Intuition

Design and implement software and machine learning components for behavior prediction and environmental interactions in a dynamic environment.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
Together AI

Join Together AI as a Staff Software Engineer to build systems that automate infrastructure management for AI clusters.

Together AI Amsterdam Published 3 weeks ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Fundamental

Join Fundamental as a Model Serving Engineer to optimize and scale the NEXUS model for enterprise decision-making.

Fundamental Europe Published 5 months ago
Flexible on stack
baseten

Join Baseten as an Infrastructure Software Engineer to build and maintain components of our ML inference platform for AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack
Pika

Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.

Pika Palo Alto HQ Published 2 months ago
Flexible on stack
baseten

Join Baseten as an AI Inference Engineer to architect and deploy high-scale production AI applications while collaborating with customers.

baseten San Francisco Published 1 month ago
Flexible on stack 70% coding
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Airbnb

Join Airbnb as a Data Scientist focusing on causal inference to enhance community support through data-driven insights.

Airbnb Remote - USA $151k–$175k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to enhance deployment infrastructure for AI systems in a collaborative environment.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 2 months ago
Flexible on stack
Instacart

Lead the design and development of core ML models for Instacart’s ads ecosystem in a fully remote role.

Instacart United States - Remote $201k–$253.5k/yr Published 3 months ago
Flexible on stack