Join Cursor as a Software Engineer on the RL Data team to create and improve tasks for training coding agents.
Posted by employer 2 weeks ago
First seen on Joblaze 2 weeks ago
Last verified on the company career page 10 hours ago
Skills & Technologies
AI in the day-to-day
We train frontier coding agents and scale RL on real user data.
Not disclosed in this posting: compensation, years of experience, work arrangement, visa sponsorship.
Joblaze summary
In the role of Software Engineer on the RL Data team at Cursor, the individual will focus on designing and refining tasks that train coding agents, ensuring they learn effectively from real user data. Key skills include strong software engineering fundamentals, experience with infrastructure or distributed systems, and an ability to analyze agent behavior to improve datasets. This position is well-suited for someone with a background in data or systems engineering, particularly those who thrive in a collaborative and innovative environment. Cursor's flat organizational structure encourages creativity and spirited discussions, making it an ideal setting for passionate problem solvers.
Joblaze insights
Quick facts
From the original posting
Our mission is to automate coding. The first step in our journey is to build the best tool for professional programmers, using a combination of inventive research, design, and engineering. Our organization is very flat, and our team is small and talent dense. We particularly like people who are truth-seeking, passionate, and creative. We enjoy spirited debate, crazy ideas, and shipping code.
Cursor is building the future of coding. We train frontier coding agents and scale RL on real user data to make them increasingly effective.
As a Software Engineer on the RL Data team at Cursor, you'll create the tasks, rewards, and environments that train our coding agents. The team owns the data that goes into training: what the model is asked to do, how we score it, and the setups it learns in.
Designing a task set that teaches a specific agent capability, then iterating on it from traces and evals until the model actually gets better.
Reading a pile of agent traces, finding a failure mode or a surprising behavior, and building a system that surfaces more of the same.
Turning a one-off recipe into something other teams can reuse: better rewards, cleaner environments, tighter data quality.
Partnering with research on whether a dataset is actually teaching the thing we think it is.
You write careful, fast code and have strong software engineering fundamentals.
You like setting tasks: breaking a fuzzy capability into something concrete you can measure.
You have an infra, data, or distributed systems background. RL experience is a plus, not a requirement.
You enjoy looking at messy real-world agent behavior and turning it into a dataset or a tool.