"large scale distributed model training" Jobs
668 open tech roles matching “large scale distributed model training”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, AI/ML. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 668 results
Join Atoms as a Staff Machine Learning Infrastructure Engineer to design and build large-scale ML training infrastructure for autonomous transport models.
Join Databricks as a Staff Software Engineer to drive the architecture of a managed GPU training platform for large-scale AI models.
Lead research and development of efficient language models at Postman, collaborating with cross-functional teams.
Operate and optimize multi-petabyte storage systems for AI workloads at Together AI.
Drive technical direction for training infrastructure and operations within Pegasus at a growing AI company focused on video understanding.
Join AI Squared as a Data Scientist to develop AI/ML solutions and collaborate with product and engineering teams in a hybrid role.
Join Skild AI as a Software Engineer to develop and optimize software infrastructure for training advanced AI models in robotics.
Join Anthropic as a Research Engineer to design and run large-scale experiments in Reinforcement Learning for AI systems.
Join DeepL as a Senior Research Scientist to innovate in reinforcement learning and shape the future of AI technology.
Join Anthropic's Inference team to design and maintain distributed systems serving AI models to millions globally.
Design and deliver multi-petabyte storage systems for AI workloads at Together AI, optimizing performance and cost.
Join Reflection AI to build systems that transform pre-trained models into aligned agents in a fast-paced startup environment.
Join Xaira Therapeutics as a Senior Software Engineer to build AI infrastructure for drug discovery and development.
Join Fundamental as an ML Researcher to tackle groundbreaking challenges in AI model development for enterprise decision-making.
Own Sarvam's production serving path for large distributed models, integrating and optimizing performance across a multi-node stack.
Related searches