"inference optimization" Jobs

214 open tech roles matching “inference optimization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 214 results

Anthropic

Join Anthropic as a Staff Software Engineer to optimize and scale AI inference across major cloud platforms.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Perplexity

Lead the financial strategy for AI products at Perplexity, optimizing model spend and driving pricing decisions.

Perplexity San Francisco Published 1 week ago
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Perplexity AI

Lead the economics of AI products at Perplexity AI, optimizing model spend and driving pricing and margin decisions.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Perplexity

Join Perplexity as a technical program manager to drive the core inference platform and coordinate between model providers and engineering teams.

Perplexity San Francisco Published 1 week ago
Anthropic

Join Anthropic as a Staff Software Engineer to design and optimize backend services for cloud inference at scale.

Anthropic San Francisco, CA $320k–$485k/yr Published 3 months ago
Flexible on stack
Chai Discovery

Join Chai Discovery as a Software Engineer to optimize AI models for drug discovery in a fast-paced, innovative environment.

Chai Discovery San Francisco office Published 9 months ago
Anthropic

Join Anthropic as a Staff Engineer to lead the technical direction of the Inference Runtime for AI systems serving millions of users.

Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | Seattle, WA | New York City, NY $405k–$485k/yr Published 3 months ago
Flexible on stack
Inferact

Join Inferact as a performance engineer to optimize vLLM, the fastest AI inference engine, working with cutting-edge hardware.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Lyft

Lead efforts in causal inference and marketing mix models to optimize marketing investments at Lyft.

Lyft San Francisco, CA $148k–$185k/yr Published 3 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as a technical program manager to drive the core inference platform and coordinate across teams and model providers.

Perplexity AI San Francisco Published 1 week ago
baseten

Join Baseten as a Software Engineer focusing on Model APIs to enhance AI model performance and developer experience.

baseten San Francisco Published 11 months ago
Faire

Own the measurement and optimization of long-term value in a hybrid role at Faire, a tech-driven wholesale platform.

Faire New York City, NY; San Francisco, CA $246.5k–$339k/yr Published 2 days ago
Flexible on stack
baseten

Join Baseten as an AI Inference Engineer to architect and deploy high-scale production AI applications while collaborating with customers.

baseten San Francisco Published 1 month ago
Flexible on stack 70% coding
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
baseten

Join Baseten as a GPU Kernel Engineer to optimize high-performance GPU kernels for cutting-edge AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack 70% coding
Inductive Bio

Join Inductive Bio as a software engineer to build AI tools that accelerate drug discovery.

Inductive Bio New York City, San Francisco, or Boston Published 5 months ago
baseten

Join Baseten as an Infrastructure Software Engineer to build and maintain components of our ML inference platform for AI applications.

baseten San Francisco Published 1 year ago
Flexible on stack