"inference systems" Jobs

404 open tech roles matching “inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 404 results

Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Anthropic

Join Anthropic's Inference team to build and maintain systems that serve AI models to millions of users worldwide.

Anthropic Ontario, CAN Published 3 weeks ago
Flexible on stack
Anthropic

Lead a team of engineers to optimize Anthropic's inference infrastructure for AI systems.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$625k/yr Published 1 week ago
Flexible on stack Heavy meetings
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 2 days ago
Flexible on stack
Anthropic

Join Anthropic as a Performance Engineer to optimize AI inference systems for throughput, latency, reliability, and correctness.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $350k–$850k/yr Published 3 months ago
Flexible on stack
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Anthropic

Join Anthropic's Inference team to design and maintain distributed systems serving AI models to millions globally.

Anthropic New York City, NY; San Francisco, CA | Seattle, WA $320k–$485k/yr Published 2 weeks ago
Flexible on stack
Anthropic

Join Anthropic's Inference team to design and maintain distributed systems that serve AI models to millions of users worldwide.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $320k–$485k/yr Published 3 months ago
Flexible on stack
Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a Site Reliability Engineer to enhance the reliability and performance of AI inference systems at scale.

Inferact San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack AI-first team
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Anthropic

Join Anthropic as a Performance Engineer to optimize the inference engine for AI systems at scale.

Anthropic San Francisco, CA | New York City, NY $350k–$850k/yr Published 3 days ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Coreweave

Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.

Coreweave Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA $188k–$275k/yr Published 6 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack