"llm infrastructure" Jobs

549 open tech roles matching “llm infrastructure”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, Kubernetes. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 549 results

Inferact

Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Blockit

Join Blockit as a Software Engineer to enhance the infrastructure of an innovative AI scheduling platform.

Blockit San Francisco Published 7 months ago
Flexible on stack
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 4 days ago
Flexible on stack
Inferact

Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco, United States Published 2 days ago
Flexible on stack
Inferact

Join Inferact as a TPU performance engineer to optimize vLLM for Google TPUs, enhancing AI inference performance.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 4 days ago
Flexible on stack
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
MaintainX

Own the LLMX roadmap as a Staff Product Manager at MaintainX, driving AI platform adoption and quality in a hybrid role.

MaintainX San Francisco Published 3 months ago
MaintainX

Own the LLMX roadmap as a Staff Product Manager at MaintainX, driving AI platform adoption and quality.

MaintainX Toronto Published 1 month ago
Braintrust Data

Join Braintrust as a backend engineer to build infrastructure for cutting-edge AI development tools in a fast-paced environment.

Braintrust Data San Francisco Published 1 month ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Sphere

Lead the development of TRAM, an AI reasoning model for interpreting global trade law in a fast-paced, onsite environment.

Sphere San Francisco HQ Published 11 months ago
Inferact
Head of Legal Hybrid Visa

Join Inferact as the first in-house legal hire to lead legal functions and support a fast-growing AI inference company.

Inferact San Francisco Published 5 days ago
Vals AI

Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
LangChain

Join LangChain as an IT Engineer to build scalable systems and support a global team in a rapidly growing environment.

LangChain San Francisco, CA $120k–$150k/yr Published 2 months ago
Flexible on stack