"inference systems" Jobs

1011 open tech roles matching “inference systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 1011 results

Coreweave

Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.

Coreweave Bellevue, WA/ San Francisco, CA/ Sunnyvale, CA $188k–$275k/yr Published 6 months ago
Flexible on stack
Pika

Join Pika as a Senior/Staff ML Engineer to enhance AI-driven products through advanced inference acceleration and GPU optimization.

Pika Palo Alto HQ Published 2 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Abridge

Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.

Abridge SF Office Published 1 year ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI UK £140k–£200k/yr Published 5 months ago
Flexible on stack
Anthropic

Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Coreweave
Coreweave Sunnyvale, CA / Bellevue, WA $165k–$242k/yr Published 11 months ago
Inworld AI

Join Inworld AI as a Staff/Principal Machine Learning Engineer to optimize and serve top-ranked realtime voice models.

Inworld AI Mountain View, California, USA $270k–$500k/yr Published 5 months ago
Flexible on stack
Chai Discovery

Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.

Chai Discovery San Francisco office Published 9 months ago
Periodic Labs

Join Periodic Labs as an ML Systems Engineer to build and optimize large-scale training and reinforcement learning infrastructure.

Periodic Labs Menlo Park, CA $250k–$350k/yr Published 4 months ago
Flexible on stack
Dialpad

Join Dialpad as a Software Engineer to build and improve ML inference systems for AI models at scale.

Dialpad Buenos Aires, Argentina Published 2 months ago
Flexible on stack 70% coding
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago
Applied Intuition

Join Applied Intuition as a Perception Software Engineer to develop safety-critical perception systems for L4 autonomous trucks.

Applied Intuition Sunnyvale Published 1 month ago
Flexible on stack
Applied Intuition

Join Applied Intuition as a Senior Software Engineer to develop perception systems for L4 autonomous trucks.

Applied Intuition Tokyo Published 2 months ago
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack
baseten

Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.

baseten San Francisco Published 5 months ago