"real time ml inference" Jobs

287 open tech roles matching “real time ml inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, PyTorch, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 287 results

Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 5 months ago
Flexible on stack
Roboflow

Join Roboflow as a Machine Learning Engineer to enhance our inference engine and contribute to impactful computer vision projects.

Roboflow NY, SF or Remote Published 2 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Pika

Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.

Pika Palo Alto HQ Published 2 months ago
Flexible on stack
Applied Intuition

Join Applied Intuition as an ML Runtime Optimization Engineer to optimize ML models for embedded environments in a collaborative team.

Applied Intuition Sunnyvale Published 1 year ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI UK £140k–£200k/yr Published 5 months ago
Flexible on stack
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Together AI

Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.

Together AI San Francisco $220k–$280k/yr Published 3 months ago
Flexible on stack 60% coding
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve state-of-the-art voice models in a fully remote role.

Inworld AI Switzerland Published 5 months ago
Flexible on stack
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Mercury

Build and operate real-time inference services for risk decisioning in a fast-growing fintech startup.

Mercury San Francisco, CA, New York, NY, Portland, OR, or Remote within Canada or United States $166.6k–$208.3k/yr Published 3 days ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 2 days ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Research Scientist to innovate in real-time voice models and AI applications.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 3 years ago
Anthropic

Lead a team of engineers to optimize Anthropic's inference infrastructure for AI systems.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$625k/yr Published 1 week ago
Flexible on stack Heavy meetings
Applied Intuition

Join Applied Intuition as an Embedded AI Engineer to develop on-device intelligence for Android Automotive platforms.

Applied Intuition Sunnyvale Published 5 months ago
Flexible on stack
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Airbnb

Lead the fine-tuning and optimization of LLMs to enhance AI products at Airbnb with a focus on customer support.

Airbnb United States $292k–$365k/yr Published 3 months ago
Flexible on stack
HappyRobot

Join HappyRobot as a Machine Learning Engineer to build AI models for human-like conversations and shape the future of AI infrastructure.

HappyRobot San Francisco Published 1 month ago
Flexible on stack
Doppel

Join Doppel as a Machine Learning Engineer to build and scale detection systems for social engineering defense.

Doppel San Francisco, New York Published 7 months ago