"low latency inference" Jobs

87 open tech roles matching “low latency inference”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 87 results

Hedra

Design and scale backend systems for Hedra’s developer platform, focusing on Python services and cloud infrastructure.

Hedra San Francisco Published 4 months ago
Databricks

Lead a multidisciplinary research team to advance large-scale machine learning efficiency at Databricks.

Databricks Mountain View, California; San Francisco, California $270k–$340k/yr Published 4 months ago
Flexible on stack 60% coding
Harvey AI

Lead the design and development of systems powering AI requests at Harvey, collaborating with multiple teams to ensure reliability and efficiency.

Harvey AI San Francisco $236k–$290k/yr Published 2 months ago
Flexible on stack
Decagon

Join Decagon as a Research Engineer to build next-generation AI voice agents in a collaborative, onsite environment.

Decagon San Francisco $200k–$400k/yr Published 1 month ago
Flexible on stack 70% coding
Fal

Join fal as a Software Engineer to build large-scale distributed systems for AI products in a growth-focused environment.

Fal San Francisco $180k–$250k/yr Published 1 year ago
Flexible on stack
Together AI

Join Together AI as a Senior Software Engineer to design and implement a scalable observability platform for our generative AI lifecycle.

Together AI San Francisco $200k–$280k/yr Published 11 months ago
Flexible on stack
Harvey

Lead the design and development of systems powering AI requests at Harvey, a fast-scaling company in the legal tech space.

Harvey San Francisco $231k–$340k/yr Published 1 month ago
Flexible on stack
baseten

Join Baseten as a Site Reliability Engineer to enhance the reliability of our multi-cloud Kubernetes infrastructure.

baseten San Francisco Published 5 months ago
Flexible on stack 70% coding
PagerDuty

Join PagerDuty as a Senior AI/ML Engineer to design and build AI-powered features for high-volume, real-time event streams.

PagerDuty Lisbon Published 1 month ago
Flexible on stack
Inductive Bio

Join Inductive Bio as a software engineer to build AI tools that accelerate drug discovery.

Inductive Bio New York City, San Francisco, or Boston Published 6 months ago
Coreweave

Join CoreWeave as a Senior Storage Engineer to design and operate high-performance file and block storage for AI workloads.

Coreweave San Francisco, CA / Sunnyvale, CA / Bellevue, WA $182k–$242k/yr Published 2 weeks ago
Flexible on stack
Atoms

Join Atoms as an ML Intern to bridge AI research and real-world applications in autonomous transport systems.

Atoms San Francisco, CA $70–$70/hr Published 5 days ago
Flexible on stack
Fireworks AI

Drive adoption of Fireworks' generative AI platform by engaging with technical founders and product teams in a fast-paced environment.

Fireworks AI San Francisco Published 5 months ago
Databricks
Databricks San Francisco, California $280k–$350k/yr Published 5 months ago
Bretton AI

Own and evolve Kubernetes infrastructure while building secure, compliant AI systems for major financial institutions.

Bretton AI San Francisco, CA $168k–$213k/yr Published 8 months ago
70% coding
Wispr Flow

Create core building blocks for AI product concepts at Wispr Flow, shaping reference architecture and shared platform primitives.

Wispr Flow San Francisco Published 1 week ago
Coreweave

Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.

Coreweave Livingston, NJ / New York, NY / Sunnyvale, CA / San Francisco, CA / Bellevue, WA $207k–$275k/yr Published 3 months ago
Flexible on stack
Pinterest

Lead the technical vision for Ads Conversion Core Modeling at Pinterest, focusing on applied ML projects and large-scale model development.

Pinterest San Francisco, CA, US; Palo Alto, CA, US; Seattle, WA, US $222.7k–$389.8k/yr Published 3 months ago