Direct from source · No middlemen

Sglang In Jobs

58 open positions · Updated 2 days ago

Average salary (USD/year): 230k–355k/yr

Showing 20 of 58 positions

Search with filters →
Inferact

Join Inferact as a co-op student to work on cutting-edge AI inference systems in a hands-on engineering role.

Inferact San Francisco Published 2 days ago
Flexible on stack
Sarvam AI

Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.

Sarvam AI Bengaluru Published 1 month ago
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 2 days ago
Flexible on stack
Inferact

Join Inferact as an AMD GPU performance engineer to optimize vLLM for the AMD accelerator ecosystem.

Inferact San Francisco $200k–$400k/yr Published 2 months ago
Flexible on stack
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems for generative AI with leading organizations.

Fireworks AI Singapore Published 1 month ago
Flexible on stack 70% coding
baseten

Join Baseten as a Solutions Architect to translate business needs into technical solutions for AI deployments.

baseten San Francisco Published 6 months ago
Inferact

Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.

Inferact Remote Published 7 months ago
Flexible on stack
ElevenLabs

Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.

ElevenLabs United Kingdom Published 2 weeks ago
Flexible on stack
Reflection AI

Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Cloudflare

Join Cloudflare as a Senior Machine Learning Engineer to optimize and productionize ML models for a global serverless inference platform.

Cloudflare Hybrid Published 2 months ago
Flexible on stack
Sarvam AI

Own the architecture of Sarvam's vision models serving harness, ensuring high-quality document intelligence at national scale.

Sarvam AI Bengaluru Published 3 weeks ago
Flexible on stack 70% coding
Fireworks AI

Join Fireworks AI as a Software Engineer to design and build scalable infrastructure for generative AI systems.

Fireworks AI San Mateo Published 10 months ago
Flexible on stack
Preference Model

Join Preference Model as a Senior Machine Learning Engineer to design RL environments for advancing ML capabilities.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Wizard

Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.

Wizard Remote - USA Published 5 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Preference Model

Join Preference Model as a senior ML Engineer to design RL environments for advancing machine learning capabilities.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Mirendil

Own the inference systems that power frontier AI models in production and research at a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
Fireworks AI

Join Fireworks AI as a senior AI Field Engineer to build production systems for innovative AI-native companies.

Fireworks AI San Mateo Published 3 months ago
Flexible on stack 70% coding