Showing 5 of 5 positions
Search with filters →Lead the engineering organization at Inferact to develop systems for vLLM, focusing on GPU performance and ML systems optimization.
Join Inferact as a cluster administration engineer to manage high-performance GPU compute infrastructure for AI inference.