"model evaluation" Jobs

4233 open tech roles matching “model evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, SQL. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 4233 results

Protege

Join Protege as a Machine Learning Researcher to lead the evaluation and optimization of audio data quality for AI training.

Protege Remote Published 3 months ago
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Arena

Join Arena as a Machine Learning Scientist to evaluate AI models and contribute to impactful research in a collaborative environment.

Arena Bay Area Published 8 months ago
Flexible on stack 70% coding
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in machine learning with a focus on RL environments.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Anthropic

Join Anthropic as a Cyber Evaluations Engineer to design and run evaluations for AI systems, ensuring their safety and robustness.

Anthropic San Francisco, CA | Washington, DC $300k–$405k/yr Published 1 week ago
Flexible on stack
Elastic

Join Elastic as an AI QA & Evaluation Engineer to validate and test AI infrastructure and solutions in a fully remote environment.

Elastic Bangalore, India Published 2 weeks ago
Flexible on stack
Reflection AI

Join Reflection AI as a hands-on technical staff member to enhance model performance through data-driven evaluations and feedback loops.

Reflection AI New York, NY Published 2 weeks ago
Reflection AI

Join Reflection AI as a Research Program Manager to build foundational infrastructure for model evaluations and safety in AI.

Reflection AI New York, NY Published 4 months ago
Ricursive Intelligence

Join Ricursive Intelligence to conduct novel AI research and work on LLM modeling and scaling in a hands-on startup environment.

Ricursive Intelligence Palo Alto Published 7 months ago
Cursor

Lead a team of engineers to build infrastructure for training and evaluating ML models in a flat, innovative organization.

Cursor San Francisco Published 2 months ago
Heavy meetings
Perplexity AI

Build specialized evals to improve answer quality across Perplexity's products in a high-impact data science role.

Perplexity AI San Francisco Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Product Designer to build and evaluate prompts for AI systems, ensuring alignment with user expectations and safety.

Anthropic San Francisco, CA $305k–$385k/yr Published 1 week ago
Flexible on stack
Reflection AI

Lead the post-training and evaluation capabilities for large language models in a dynamic AI research lab.

Reflection AI New York, NY Published 10 months ago
Periodic Labs

Join Periodic Labs as a Midtraining Research Engineer to enhance scientific reasoning in AI models for groundbreaking discoveries.

Periodic Labs Menlo Park, CA $250k–$350k/yr Published 1 month ago
Anthropic

Drive model launches and improve coding performance as a Product Manager on Claude Code's model performance team.

Anthropic San Francisco, CA | Seattle, WA $305k–$460k/yr Published 3 months ago
Cartesia

Join Cartesia as a Researcher to enhance multimodal models through innovative post-training methods and alignment techniques.

Cartesia *HQ - San Francisco, CA Published 10 months ago
Figma

Join Figma as a Design Program Manager to enhance AI evaluation processes and drive design quality.

Figma London, England Published 3 days ago
Cartesia

Join Cartesia as a Researcher to advance neural network architecture design in a collaborative, innovative environment.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack