"real time inference" Jobs

526 open tech roles matching “real time inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 526 results

Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 5 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI UK £140k–£200k/yr Published 5 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve state-of-the-art voice models in a fully remote role.

Inworld AI Switzerland Published 5 months ago
Flexible on stack
Anthropic

Join Anthropic's Inference team to build and maintain systems that serve AI models to millions of users worldwide.

Anthropic Ontario, CAN Published 3 weeks ago
Flexible on stack
Coreweave

Lead complex, cross-functional programs for inference platform delivery at a rapidly growing AI cloud company.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / Bellevue, WA $198k–$264k/yr Published 2 months ago
Pika

Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.

Pika Palo Alto HQ Published 2 months ago
Flexible on stack
Roboflow

Join Roboflow as a Machine Learning Engineer to enhance our inference engine and contribute to impactful computer vision projects.

Roboflow NY, SF or Remote Published 2 months ago
Flexible on stack
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Coreweave

Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.

Coreweave Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA $188k–$275k/yr Published 6 months ago
Flexible on stack
Coreweave

Join CoreWeave as a Staff Software Engineer to lead the development of a Kubernetes-native inference platform for AI workloads.

Coreweave Sunnyvale, CA / Bellevue, WA $188k–$275k/yr Published 4 months ago
Flexible on stack
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 3 months ago
Flexible on stack
Anthropic

Lead a team of engineers to optimize Anthropic's inference infrastructure for AI systems.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$625k/yr Published 1 week ago
Flexible on stack Heavy meetings
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Applied Intuition

Design and implement software and machine learning components for behavior prediction and environmental interactions in a dynamic environment.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
Anthropic

Join Anthropic's Inference team to design and maintain distributed systems that serve AI models to millions of users worldwide.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 3 months ago
Flexible on stack
Inflection AI

Lead the development of Inflection's realtime Voice AI stack, shaping emotionally intelligent AI for enterprise voice interactions.

Inflection AI Palo Alto, California, United States $400k–$550k/yr Published 2 months ago
Applied Intuition

Join Applied Intuition as an AI Performance Engineer to optimize large-scale machine learning workloads in a collaborative environment.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack