Direct from source · No middlemen
33 open positions · Updated 1 week ago
33 Sglang roles across 18 companies, most in AI/ML; 9% fully remote; typical advertised salary $196k
Who is hiring (18 companies)
Role types
Work arrangement: 3 fully remote · 4 hybrid · 9 on-site · 17 not stated
Advertised salaries: p25 $181k · median $196k · p75 $205k (from 16 disclosed annual salaries, USD)
Counts are open roles Joblaze currently tracks on company career pages for this exact skill/location; salaries are advertised minimums, annual, converted to USD.
Showing 20 of 33 positions
Search with filters →Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.
Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.
Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.
Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.
Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.
Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.
Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.
Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Join Periodic Labs as an ML Systems Engineer to build and optimize large-scale training and reinforcement learning infrastructure.
Own the inference systems that power frontier AI models in production and research at a tech-first startup.
Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.
Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.
Join SpaceX as a Software Engineer to develop high-performance AI inference systems for mission-critical applications.
Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.