"inference infrastructure" Jobs

934 open tech roles matching “inference infrastructure”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 934 results

Fireworks AI

Join Fireworks AI as an AI Field Engineer to build production systems for generative AI with large organizations across EMEA.

Fireworks AI London Published 1 month ago
Flexible on stack 70% coding
Sarvam AI

Join Sarvam as an Infrastructure SRE to operate a large GPU fleet and solve complex reliability challenges in AI workloads.

Sarvam AI Bengaluru Published 2 months ago
Flexible on stack
Mercury

Build and operate real-time inference services for risk decisioning in a fast-growing fintech startup.

Mercury San Francisco, CA, New York, NY, Portland, OR, or Remote within Canada or United States $166.6k–$208.3k/yr Published 4 days ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Reflection AI

Lead the post-training and evaluation capabilities for large language models in a dynamic AI research lab.

Reflection AI New York, NY Published 10 months ago
Anthropic

Join Anthropic as a Regional Research Economist to analyze AI's economic impact and shape policy discussions.

Anthropic London, UK £180k–£190k/yr Published 3 months ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Regional Research Economist to analyze AI's economic impact and shape policy discussions.

Anthropic Singapore S$307.2k–S$331.2k/yr Published 3 months ago
Flexible on stack
Inworld AI

Join Inworld AI as a Staff/Principal Platform Engineer to build and scale AI products with a focus on cloud infrastructure.

Inworld AI Mountain View, California, USA $280k–$350k/yr Published 6 months ago
Flexible on stack
Together AI

Operate and optimize multi-petabyte storage systems for AI workloads at Together AI.

Together AI Bangalore India Published 2 days ago
Flexible on stack
Twelve Labs

Lead and build a new team focused on developing Jockey Core, a reasoning LLM for video understanding at Twelve Labs.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
OpenRouter

Join OpenRouter as a Forward Deployed Engineer to help customers implement and scale AI solutions effectively.

OpenRouter San Francisco Bay Area, California Published 3 months ago
Flexible on stack 70% coding
Fal

Own the reliability and security of fal's generative media model APIs in a hybrid ML Engineering/SRE role.

Fal Remote - APAC Published 2 months ago
Flexible on stack
Poolside AI

Join Poolside AI's compute team to optimize GPU utilization and enhance inference serving for cutting-edge AI research.

Poolside AI Remote (EMEA) Published 2 months ago
Flexible on stack
Doppel

Join Doppel as a Machine Learning Engineer to build and scale detection systems against evolving digital threats.

Doppel Toronto, ON Published 4 months ago
Sarvam AI

Own the architecture of Sarvam's vision models serving harness, ensuring high-quality document intelligence at national scale.

Sarvam AI Bengaluru Published 3 weeks ago
Flexible on stack 70% coding
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack