52 jobs tagged vLLM
Join Inferact as an inference runtime engineer to innovate AI inference engines for large models in a fully remote role.
Join ChipAgents as an ML Systems Engineer to optimize LLM inference systems for leading semiconductor companies.
Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.
Own the model lifecycle for defence and strategic sector deployments as an MLOps Engineer at Sarvam AI.
Own the end-to-end lifecycle of production ML serving systems for a top-performing AI Shopping Agent.
Join Handshake as a Senior Software Engineer to build scalable ML infrastructure for a fast-growing AI data business.
Join Baseten as a Software Engineer to build and optimize large-scale LLM inference systems in a collaborative environment.
Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.
Join Inferact as a staff engineer to work on optimizing AI inference across the vLLM stack in a fully remote role.
Join Inferact as a Developer Relations Engineer to shape how developers learn and build with vLLM, the AI inference engine.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Join Inferact as an inference runtime engineer to optimize AI model execution across diverse hardware and architectures.
Own the inference systems that power frontier AI models in production and research at a tech-first startup.
Join Sola as a Software Engineer to build reliable AI features in a fast-growing startup environment in NYC.
Join Sesame as an ML Model Serving Engineer to enhance our serving layer for voice agents with cutting-edge techniques.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join Stuut as a Member of the Technical Staff to design and deploy AI-powered systems for financial operations.
Join Abridge as a Machine Learning Infrastructure Engineer to optimize AI model inference infrastructure in a fast-paced healthcare startup.
Drive the adoption of AI runtime services at CoreWeave, leveraging your expertise in distributed systems and AI infrastructure.
Join CoreWeave as a Staff Software Engineer to lead the development of a Kubernetes-native inference platform for AI workloads.
Join CoreWeave as an Applied AI Engineer to enhance the performance of our inference platform through benchmarking and optimization.
Join ElevenLabs as a Research Engineer to deploy and optimize AI models for real-time applications in a fully remote environment.
Join SpaceX as a Software Engineer to develop high-performance AI inference systems for mission-critical applications.
Join Together AI as a Forward Deployed Engineer to optimize inference systems for strategic customers in a hands-on role.
Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.
Join Fireworks AI as a senior AI Field Engineer to build production systems for generative AI with leading organizations.
Join Fireworks AI as an AI Field Engineer to build production systems for generative AI with large organizations across EMEA.
Join Fireworks AI as a senior AI Field Engineer to build production systems and engage with enterprise customers on generative AI solutions.
Join Fireworks AI as a senior AI Field Engineer to build production systems for innovative AI-native companies.
Join Cloudflare as a Senior Machine Learning Engineer to optimize and productionize ML models for a global serverless inference platform.
Join Together AI as a Research Engineer to develop a platform for customizing open-source models with user data.
Join Roboflow as a Machine Learning Engineer to enhance our inference engine and contribute to impactful computer vision projects.
Join Together AI as a Staff ML Engineer to optimize voice model serving for real-time applications on a high-impact team.