← Back to results

Software Engineer, Inference

Join Sierra as a Software Engineer on the Inference team to build efficient AI systems for customer-facing applications.

Location
San Francisco, CA, United States
Compensation
Not disclosed
Level
mid
Type
full time · On-site

Posted by employer 1 day ago

First seen on Joblaze 1 hour ago

Last verified on the company career page 1 hour ago

Apply at Sierra → Save job Scanned from sierra.ai

What you'll build

  • Define Sierra’s inference architecture
  • Develop systems for routing and failover
  • Run models on GPU infrastructure
  • Optimize inference performance
  • Contribute to infrastructure that enables post-training

Must have

  • Deep systems thinking
  • Experience designing large-scale production systems
  • Strong judgment around tradeoffs
  • Experience taking ownership of complex infrastructure

Nice to have

  • Experience with ML infrastructure
  • Experience serving LLMs at scale
  • Experience operating self-hosted inference
  • Familiarity with inference frameworks
  • Experience with post-training infrastructure

AI in the day-to-day

Sierra’s AI agents depend on foundation models to reason and act in real time.

Not disclosed in this posting: compensation, years of experience, visa sponsorship.

Benefits

Retirement plan dependent on country of employment Discretionary benefit stipend Medical, Dental, and Vision benefits Flexible (unlimited) paid time off Lunch and Snacks Fertility and family building benefits Life Insurance and Disability Benefits Parental Leave

Joblaze summary

In the role of Software Engineer on the Inference team at Sierra, the individual will focus on designing and optimizing systems that ensure the efficient and reliable operation of AI models in real-time. Key skills include a strong foundation in distributed systems and experience with large-scale production environments, particularly in managing latency and capacity. This position is well-suited for engineers with a background in complex systems who are eager to apply their expertise to AI infrastructure. Sierra emphasizes a collaborative culture, driven by values such as trust and customer obsession.

Joblaze insights

  • Listed today — first seen on Joblaze October 7, 2026. Last confirmed on Sierra's careers page October 7, 2026.
  • AI/ML appears in 52.2% of 462 comparable mid ai/ml roles in United States; infrastructure appears in 0.9% of 462 comparable mid ai/ml roles in United States.

Quick facts

Is the Software Engineer, Inference role remote?
No — this is an on-site role in San Francisco, CA, United States.
Where is the role based?
Sierra is hiring for this position in San Francisco, CA, United States.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, GPU, MLOps, infrastructure.
What seniority level is this role?
Sierra targets mid-level candidates for this position.
Is this full-time or contract?
Full-time for this Software Engineer, Inference role at Sierra.

From the original posting

About the role

Sierra’s AI agents depend on foundation models to reason and act in real time. The Inference team builds the systems that make those models fast, reliable, and efficient at scale.

As a Software Engineer on Inference, you’ll help define Sierra’s inference architecture across both self-hosted models and third-party inference providers. You’ll work on the systems responsible for serving and routing inference, managing capacity and quota, and optimizing for latency, reliability, and cost.

This is a systems-first role at the intersection of distributed infrastructure and AI. You don’t need to be an ML researcher—we’re looking for engineers who love complex systems problems and are excited to apply that expertise to one of the fastest-moving areas of AI infrastructure.

What you'll do

  • Partner with frontier labs and providers. At our scale, we rely on frontier labs, and inference providers to supply capacity, training and inference infrastructure.

  • Shape Sierra’s inference architecture. Design how inference traffic flows across models, infrastructure, and providers, including new serving and proxy layers as Sierra scales.

  • Build for low latency and high reliability. Develop systems for routing, failover, capacity management, and quota that keep inference performant and available across large-scale production workloads.

  • Build and operate self-hosted inference. Run models on GPU infrastructure, from building containers and operating inference engines to managing the underlying compute capacity.

  • Optimize inference performance. Work with the Applied Research team on techniques such as speculative decoding and serving-engine optimizations that improve latency, throughput, and cost.

  • Build across a hybrid inference stack. Work with both Sierra-managed infrastructure and leading inference platforms, making architectural decisions about where and how workloads should run.

  • Push the serving stack forward. Work closely with inference providers to tune engines and infrastructure for Sierra’s workloads.

  • Support the broader model lifecycle. Contribute to infrastructure that enables post-training while partnering closely with our Models and Agent Runtime teams.

What you'll bring

  • Deep systems thinking and strong distributed systems fundamentals.

  • Experience designing, building, and operating large-scale production systems.

  • Strong judgment around tradeoffs involving latency, reliability, capacity, and cost.

  • Experience taking ownership of complex infrastructure from architecture through production operation.

  • Excitement about applying systems expertise to AI infrastructure and learning quickly as the underlying technology evolves.

Even better

  • Experience with ML infrastructure, MLOps, or production inference systems.

  • Experience serving LLMs or other large models at scale.

  • Experience operating self-hosted inference and GPU infrastructure.

  • Familiarity with inference frameworks such as vLLM or SGLang.

  • Experience with post-training infrastructure or inference-performance optimization.

Our values

  • Trust: We build trust with our customers with our accountability, empathy, quality, and responsiveness. We build trust in AI by making it more accessible, safe, and useful. We build trust with each other by showing up for each other professionally and personally, creating an environment that enables all of us to do our best work.

  • Flexible (unlimited) paid time off

  • Life insurance and disability benefits

  • Retirement plan dependent on country of employment

  • Parental leave

  • Fertility and family building benefits through Carrot

  • Free alphorn lessons

Standard company text repeated across Sierra's postings is omitted here.

Similar positions

Sierra
Software Engineer, Infrastructure
Sierra · San Francisco, CA
Sierra
Software Engineer, Insights
Sierra · San Francisco, CA
Sierra
Research Engineer, Applied Research
Sierra · San Francisco, CA, United States
Sierra
Software Engineer, Context Engine
Sierra · San Francisco, CA
Sierra
Software Engineer, Platform
Sierra · San Francisco, CA