"low latency inference" Jobs

9 open tech roles matching “low latency inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, GPU. Every listing is re-checked daily and closed roles are removed.

Showing 9 of 9 results

Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 2 months ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 2 years ago
Mercury

Join Mercury as a Senior Machine Learning Operations Engineer to build and operate real-time inference services for risk decisioning.

Mercury San Francisco, CA, New York, NY, Portland, OR, or Remote within Canada or United States $166.6k–$208.3k/yr Published 2 months ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $280k–$850k/yr Published 1 year ago
Fireworks AI

Design and maintain large-scale backend infrastructure for a leading generative AI platform at Fireworks AI.

Fireworks AI New York Published 2 months ago
Flexible on stack
Fireworks AI

Own foundational capabilities for enterprise AI, designing data models and APIs while ensuring security and compliance for large-scale customers.

Fireworks AI New York Published 2 weeks ago
Flexible on stack
Fireworks AI

Drive sourced pipeline and revenue through the Microsoft Azure channel in a high-ownership sales role at Fireworks AI.

Fireworks AI New York Published 1 month ago