"rl training" Jobs

220 open tech roles matching “rl training”, taken straight from company career pages — not reposted from other job boards. Every listing is re-checked daily and closed roles are removed.

No "rl training" jobs in Berlin right now — showing "rl training" jobs in all locations.

Showing 20 of 220 results

Mirendil

Join Mirendil as a research engineer to build the post-training stack for frontier reasoning models in AI.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in large language models within a fast-paced startup.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in machine learning with a focus on RL environments.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Reflection AI

Lead the post-training and evaluation capabilities for large language models in a dynamic AI research lab.

Reflection AI New York, NY Published 10 months ago
Periodic Labs

Join Periodic Labs as a Midtraining Research Engineer to enhance scientific reasoning in AI models for groundbreaking discoveries.

Periodic Labs Menlo Park, CA $250k–$350k/yr Published 1 month ago
Reflection AI

Join Reflection AI to build systems that transform pre-trained models into aligned agents in a fast-paced startup environment.

Reflection AI San Francisco, CA Published 1 month ago
Cursor

Join Cursor as a Software Engineer on the RL Data team to create and improve tasks for training coding agents.

Cursor San Francisco Published 2 weeks ago
baseten

Join Baseten as a senior software engineer to develop cutting-edge AI training products and enhance user workflows.

baseten San Francisco Published 7 months ago
Flexible on stack
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Mirendil

Join Mirendil as a staff engineer to build the post-training stack for frontier reasoning models in a tech-first startup.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
Harvey AI

Join Harvey AI as a Research Engineer to drive post-training experiments and enhance legal AI models.

Harvey AI San Francisco $231k–$340k/yr Published 2 months ago
Flexible on stack
Distyl AI

Join Distyl AI as a Senior Applied AI Researcher to redefine AI utilization in enterprise with cutting-edge research and technology.

Distyl AI San Francisco $150k–$250k/yr Published 11 months ago
Flexible on stack
Hippocratic AI

Join Hippocratic AI as an Applied Scientist to enhance AI safety and clinical reasoning through reinforcement learning.

Hippocratic AI Menlo Park, CA Published 1 month ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 2 days ago
Flexible on stack
Reflection AI

Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.

Reflection AI San Francisco, CA Published 8 months ago
Flexible on stack
Reflection AI

Join Reflection AI as a Research Software Engineer to bridge research and production in cutting-edge AI training systems.

Reflection AI New York, NY Published 6 months ago
Flexible on stack
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
DeepL

Join DeepL as a Senior Research Scientist to innovate in reinforcement learning and shape the future of AI technology.

DeepL London Published 1 month ago
Flexible on stack