"large scale distributed model training" Jobs

154 open tech roles matching “large scale distributed model training”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, Go. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 154 results

Reflection AI

Build and scale distributed training systems for frontier model pre-training at Reflection AI.

Reflection AI San Francisco, CA Published 5 months ago
Flexible on stack
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in large language models within a fast-paced startup.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in machine learning with a focus on RL environments.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI inference workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Fireworks AI

Design and maintain large-scale backend infrastructure for a leading generative AI platform at Fireworks AI.

Fireworks AI New York Published 3 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact San Francisco $200k–$400k/yr Published 7 months ago
Flexible on stack
Inferact

Join Inferact as a staff engineer to build distributed systems for AI inference at global scale.

Inferact Singapore S$200k–S$400k/yr Published 2 months ago
Flexible on stack
Mirendil

Join Mirendil as a staff engineer to work on cutting-edge AI pretraining, optimizing models and infrastructure.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to design and optimize large-scale AI training and inference clusters.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to build LLM infrastructure for large-scale AI workloads.

Databricks San Francisco, California $190k–$265k/yr Published 1 month ago
Flexible on stack
Cantina

Join Cantina as a Member of Technical Staff to build and scale data pipelines for large video generation models.

Cantina Remote (U.S. or Europe) $200k–$260k/yr Published 5 months ago
Flexible on stack
Reflection AI

Join Reflection AI as a Research Software Engineer to bridge research and production in cutting-edge AI training systems.

Reflection AI New York, NY Published 6 months ago
Flexible on stack
Perplexity AI

Join Perplexity AI as an AI Infrastructure Engineer to build and optimize large-scale AI training and inference clusters.

Perplexity AI London Published 5 months ago
Flexible on stack
Reddit

Lead the development of large-scale machine learning infrastructure to enhance personalization and recommendation systems at Reddit.

Reddit Remote - United States $253.3k–$354.6k/yr Published 1 month ago
Flexible on stack
Atoms

Join Atoms as a Staff Machine Learning Infrastructure Engineer to design and build large-scale ML training infrastructure for autonomous transport models.

Atoms San Francisco, CA $224k–$280k/yr Published 2 months ago
Flexible on stack
Databricks

Join Databricks as a Staff Software Engineer to drive the architecture of a managed GPU training platform for large-scale AI models.

Databricks Mountain View, California; San Francisco, California $190k–$265k/yr Published 3 months ago