"large scale distributed model training" Jobs

260 open tech roles matching “large scale distributed model training”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, Kubernetes, PyTorch. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 260 results

Cartesia

Join Cartesia as a Product and Research Operations Manager to design and scale a global evaluation workforce for AI.

Cartesia *HQ - San Francisco, CA Published 2 weeks ago
AI-first team
Goodfire

Join Goodfire as a Machine Learning Engineer to build interpretable AI systems with a world-class team.

Goodfire San Francisco, CA & New York, NY $200k–$400k/yr Published 9 months ago
Graphcore

Join Graphcore as a Senior Machine Learning Engineer to advance AI technology on cutting-edge hardware.

Graphcore Gdańsk, Pomeranian Voivodeship, Poland PLN 260.4k–PLN 352.2k/yr Published 2 months ago
Flexible on stack
Protege

Join Protege as a Machine Learning Researcher to lead the evaluation and optimization of audio data quality for AI training.

Protege Remote Published 3 months ago
Fundamental

Join Fundamental as a senior software engineer to enhance ML model development and infrastructure in a pioneering AI startup.

Fundamental Barcelona Published 4 months ago
Cantina

Join Cantina as a Machine Learning Engineer to build advanced speech systems and contribute to innovative AI technology.

Cantina Europe $200k–$220k/yr Published 4 months ago
Flexible on stack 70% coding
Affirm

Lead the development of fraud prediction models in a remote-first team at Affirm.

Affirm Remote Canada CA$153k–CA$213k/yr Published 1 month ago
Flexible on stack
Databricks

Join Databricks as a Senior Applied ML Engineer to optimize infrastructure and enhance serverless compute products.

Databricks San Francisco, California $16k–$21k/mo Published 1 month ago
Flexible on stack
Cantina

Join Cantina as a Machine Learning Engineer to develop cutting-edge speech and audio generation systems in a collaborative environment.

Cantina Remote (U.S. or Europe) $200k–$220k/yr Published 1 month ago
Flexible on stack 70% coding
Sarvam AI

Join Sarvam as an Infrastructure SRE to operate a large GPU fleet and solve complex reliability challenges in AI workloads.

Sarvam AI Bengaluru Published 2 months ago
Flexible on stack
Hippocratic AI

Own the serving infrastructure for healthcare AI, optimizing LLM inference systems to enhance patient experiences.

Hippocratic AI Menlo Park, CA Published 3 weeks ago
Flexible on stack
Coreweave

Join CoreWeave as a Senior Applied ML Engineer to tackle challenges in continuous learning for AI agents with a focus on innovative solutions.

Coreweave Bellevue, WA / Sunnyvale, CA $182k–$242k/yr Published 5 months ago
Flexible on stack
Together AI

Build production AI agents and foundational systems for one of the world's largest GPU fleets at Together AI in Amsterdam.

Together AI Amsterdam Published 1 week ago
Flexible on stack
Together AI

Build production AI agent systems for one of the world's largest GPU fleets at Together AI.

Together AI San Francisco $250k–$300k/yr Published 4 days ago
Flexible on stack
Cartesia

Join Cartesia as an Inference Engineer to design and build low latency, scalable model inference for cutting-edge AI applications.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack