"inference optimization" Jobs

209 open tech roles matching “inference optimization”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 209 results

Attentive Mobile

Lead a high-performing data science team to drive product strategy and business growth through innovative analytical solutions.

Attentive Mobile San Francisco, CA $320k–$380k/yr Published 2 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a Senior Field Marketing Manager to create innovative events that shape the brand experience on the West or East Coast.

Fireworks AI San Francisco Published 3 months ago
Anthropic

Lead the Capacity Engineering team at Anthropic, ensuring efficient allocation and utilization of infrastructure resources.

Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$485k/yr Published 3 days ago
Flexible on stack Heavy meetings
baseten

Lead the capacity management for Baseten's TPU fleet, ensuring optimal performance and reliability in AI workloads.

baseten San Francisco Published 4 days ago
Flexible on stack
Databricks
Databricks Mountain View, California; San Francisco, California Published 6 months ago
Amplitude

Lead a multidisciplinary team to advance experimentation methodologies and data infrastructure at Amplitude.

Amplitude San Francisco, CA $227k–$381k/yr Published 1 month ago
Heavy meetings
Asana

Lead Asana's AI Platform organization to drive strategy and execution for AI experiences across the company.

Asana San Francisco $306k–$360k/yr Published 1 month ago
Amplitude

Lead a multidisciplinary team to enhance experimentation data infrastructure at Amplitude, a leading AI analytics platform.

Amplitude San Francisco, CA $227k–$381k/yr Published 1 month ago
Flexible on stack Heavy meetings
PagerDuty

Join PagerDuty as a Senior AI/ML Engineer to design and build AI-powered features for high-volume, real-time event streams.

PagerDuty Lisbon Published 3 weeks ago
Flexible on stack
Decagon

Join Decagon as a Research Engineer to build next-generation AI voice agents in a collaborative, onsite environment.

Decagon San Francisco $200k–$400k/yr Published 1 week ago
Flexible on stack 70% coding
Anthropic
Anthropic San Francisco, CA | New York City, NY | Seattle, WA $405k–$485k/yr Published 5 months ago
Coreweave

Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA $207k–$275k/yr Published 2 months ago
Flexible on stack
Fal

Join fal as a Senior Solutions Architect to engage with clients and design scalable AI solutions in a growing generative media ecosystem.

Fal San Francisco $220k–$290k/yr Published 1 month ago
Flexible on stack
Decagon

Design and operate data systems that power Decagon's AI products, ensuring high reliability and performance.

Decagon San Francisco $200k–$400k/yr Published 3 weeks ago
Flexible on stack
Harvey AI

Lead a high-performing team to build and operate Harvey's core compute and networking infrastructure for AI-driven legal services.

Harvey AI San Francisco $272k–$355k/yr Published 2 months ago
Flexible on stack Heavy meetings
Postman

Join Postman as an AI Engineer Intern to work on large-scale AI systems with mentorship from senior engineers.

Postman Berkeley, California, United States; San Francisco, California, United States Published 1 month ago
Flexible on stack
Decagon

Lead the development of models and algorithms for Decagon's real-time voice agents in a collaborative, onsite environment.

Decagon San Francisco $200k–$400k/yr Published 3 months ago
Flexible on stack 70% coding
Gamma

Build and scale backend systems for millions of users at Gamma, focusing on performance and reliability in a collaborative environment.

Gamma San Francisco $180k–$310k/yr Published 3 months ago
Flexible on stack
Wispr Flow
ML Engineer Hybrid Visa

Join Wispr Flow as a ML Engineer to build a scalable voice interface for millions of users.

Wispr Flow San Francisco Published 1 year ago
Flexible on stack