"inference infrastructure" Jobs

402 open tech roles matching “inference infrastructure”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 402 results

Twelve Labs

Build and operate production ML systems for Pegasus, focusing on reliability and performance in a hybrid work environment.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
Goodfire

Join Goodfire as a Product Engineer to build user-friendly AI systems with a focus on interpretability and product development.

Goodfire San Francisco, CA Published 3 months ago
baseten

Join Baseten as a Software Engineer to build continuous deployment infrastructure for AI products in a fast-growing team.

baseten San Francisco Published 1 month ago
Flexible on stack
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 2 weeks ago
Flexible on stack
Anthropic
Anthropic New York City, NY; San Francisco, CA; Seattle, WA $275k–$370k/yr Published 4 months ago
baseten

Join Baseten to lead capacity planning and operations in a fast-paced AI environment.

baseten San Francisco Published 1 week ago
Harvey AI

Join Harvey as a Staff Product Manager to shape the infrastructure of a leading legal AI platform with a focus on reliability and scalability.

Harvey AI San Francisco $213.6k–$300k/yr Published 2 months ago
AI-first team
Mercury

Build and operate real-time inference services for risk decisioning in a fast-growing fintech startup.

Mercury San Francisco, CA, New York, NY, Portland, OR, or Remote within Canada or United States $166.6k–$208.3k/yr Published 1 week ago
Flexible on stack
baseten

Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.

baseten San Francisco Published 4 months ago
Flexible on stack Heavy meetings
Twelve Labs

Lead and build a new team focused on developing Jockey Core, a reasoning LLM for video understanding at Twelve Labs.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
OpenRouter

Join OpenRouter as a Forward Deployed Engineer to help customers implement and scale AI solutions effectively.

OpenRouter San Francisco Bay Area, California Published 3 months ago
Flexible on stack 70% coding
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 3 months ago
Flexible on stack
baseten

Join Baseten as a Sr. Analyst in Revenue Strategy & Operations to shape GTM strategies for AI infrastructure.

baseten San Francisco Published 2 weeks ago
baseten

Join Baseten as a Forward Deployed Engineer to architect and deploy high-scale AI applications while collaborating with customers.

baseten San Francisco Published 2 years ago
Flexible on stack 70% coding
Together AI

Join Together AI as a Senior Software Engineer to design and implement a scalable observability platform for our generative AI lifecycle.

Together AI San Francisco $200k–$280k/yr Published 10 months ago
Flexible on stack
Insitro

Join insitro as a Senior ML Scientist to develop machine learning methods for biological data analysis in a vibrant biotech startup.

Insitro South San Francisco, CA $183k–$238k/yr Published 4 months ago
Flexible on stack
baseten

Join Baseten as a Strategic Finance Associate to support financial planning and analysis in a fast-scaling AI infrastructure company.

baseten San Francisco Published 3 months ago