"rl training" Jobs
110 open tech roles matching “rl training”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, reinforcement learning, AI/ML. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 110 results
Join Mirendil as a research engineer to build the post-training stack for frontier reasoning models in AI.
Join Preference Model as a Research Engineer to advance self-directed learning in large language models within a fast-paced startup.
Join Preference Model as a Research Engineer to advance self-directed learning in machine learning with a focus on RL environments.
Join Reflection AI to build systems that transform pre-trained models into aligned agents in a fast-paced startup environment.
Join Cursor as a Software Engineer on the RL Data team to create and improve tasks for training coding agents.
Join Baseten as a senior software engineer to develop cutting-edge AI training products and enhance user workflows.
Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.
Join Mirendil as a staff engineer to build the post-training stack for frontier reasoning models in a tech-first startup.
Join Harvey AI as a Research Engineer to drive post-training experiments and enhance legal AI models.
Join Distyl AI as a Senior Applied AI Researcher to redefine AI utilization in enterprise with cutting-edge research and technology.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.
Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.
Join Anthropic as a Technical Program Manager to drive progress in reinforcement learning research and enhance AI systems.
Join Anthropic as a Research Engineer to enhance AI's coding capabilities through reinforcement learning in a collaborative environment.
Design and operate large-scale GPU infrastructure for model inference and mid-training workloads at Reflection AI.
Join Labelbox as a Staff ML Engineer to shape AI training environments and systems in a high-impact, fast-paced setting.