"ai model evaluation" Jobs

3713 open tech roles matching “ai model evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, SQL. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 3713 results

krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Harvey AI

Join Harvey as a Senior Product Operations Manager to build and scale the evaluation engine for a global AI platform.

Harvey AI San Francisco Published 3 months ago
AI-first team
Anthropic

Drive model launches and improve coding performance as a Product Manager on Claude Code's model performance team.

Anthropic San Francisco, CA | Seattle, WA $305k–$460k/yr Published 3 months ago
Sarvam AI

Join Sarvam AI as a Product Manager for Model APIs, influencing the model roadmap and developer experience.

Sarvam AI Bengaluru Published 2 weeks ago
AI-first team
Inflection AI

Lead model training and post-training strategies for emotionally intelligent AI at Inflection AI.

Inflection AI Palo Alto, California, United States $400k–$550k/yr Published 2 months ago
Cursor

Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.

Cursor San Francisco Published 3 months ago
Heavy meetings
Protege

Join Protege as a Machine Learning Researcher to lead the evaluation and optimization of audio data quality for AI training.

Protege Remote Published 3 months ago
Cartesia

Join Cartesia as a Researcher to advance neural network architecture design in a collaborative, innovative environment.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Benchling

Join Benchling as a Research Engineer to improve AI models for scientific applications in a fast-paced, collaborative environment.

Benchling San Francisco, CA Published 2 weeks ago
OpenRouter

Conduct original research on large language models to advance understanding and routing optimization at OpenRouter.

OpenRouter Remote (US) Published 2 months ago
Flexible on stack
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in large language models within a fast-paced startup.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Anthropic

Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.

Anthropic San Francisco, CA $500k–$850k/yr Published 1 month ago
Flexible on stack
Macroscope

Join Macroscope as an Applied ML Engineer to enhance machine learning systems in a collaborative startup environment.

Macroscope San Francisco $170k–$280k/yr Published 1 month ago
Flexible on stack
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in machine learning with a focus on RL environments.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Reflection AI

Join Reflection AI as a hands-on technical staff member to enhance model performance through data-driven evaluations and feedback loops.

Reflection AI New York, NY Published 2 weeks ago
Cartesia

Join Cartesia as a Researcher in London to advance AI through innovative neural network architecture design.

Cartesia London Published 11 months ago
Flexible on stack
Reflection AI

Join Reflection AI as a Research Program Manager to build foundational infrastructure for model evaluations and safety in AI.

Reflection AI New York, NY Published 4 months ago
Anthropic
Anthropic San Francisco, CA | New York City, NY | Washington, DC $230k–$270k/yr Published 6 months ago