"sglang" Jobs

31 open tech roles matching “sglang”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: vLLM, SGLang, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 11 of 31 results

Preference Model

Join Preference Model as a senior ML Engineer to design RL environments for advancing machine learning capabilities.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.

baseten San Francisco Published 3 months ago
Flexible on stack
Inferact

Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.

Inferact San Francisco Published 1 month ago
baseten

Join Baseten as a Forward Deployed Engineer to solve complex AI challenges for leading companies.

baseten San Francisco Published 3 weeks ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to lead GPU Networking efforts and optimize distributed systems for AI applications.

baseten San Francisco Published 6 months ago
Flexible on stack
Together AI

Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.

Together AI San Francisco $200k–$290k/yr Published 2 months ago
Flexible on stack
Twelve Labs

Lead and build a new team focused on developing Jockey Core, a reasoning LLM for video understanding at Twelve Labs.

Twelve Labs Seoul, South Korea Published 1 month ago
Flexible on stack
baseten

Join Baseten as a Technical Program Manager to build and optimize the core algorithms for high-performance AI inference.

baseten San Francisco Published 2 weeks ago
baseten

Join Baseten as a Product Manager to shape the future of AI infrastructure and enhance production inference capabilities.

baseten San Francisco Published 5 months ago
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack