"model evaluation" Jobs

4234 open tech roles matching “model evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, SQL. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 4234 results

Cartesia

Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.

Cartesia *HQ - San Francisco, CA Published 10 months ago
AI-first team
Triomics

Join Triomics as an ML Evaluation Engineer to ensure model quality and stability in clinical AI systems.

Triomics India Office Published 2 months ago
Flexible on stack
Cantina

Join Cantina as a Machine Learning Engineer to lead audio model evaluations and shape the future of social AI.

Cantina Europe $200k–$220k/yr Published 4 months ago
AI-first team
Mirendil

Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
AI-first team
Sarvam AI

Join Sarvam as a Data Scientist to design evaluation frameworks for AI outputs in high-stakes domains.

Sarvam AI Bengaluru Published 4 months ago
Flexible on stack
Exa

Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.

Exa San Francisco, California Published 11 months ago
Flexible on stack
Sesame

Join Sesame as a Research Engineer to innovate in NLP, Speech, and Computer Vision with a focus on deep learning.

Sesame San Francisco Published 2 months ago
Flexible on stack
Anthropic
Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY $320k–$485k/yr Published 4 months ago
Reflection AI

Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.

Reflection AI San Francisco, CA Published 8 months ago
Cursor

Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.

Cursor San Francisco Published 3 months ago
Heavy meetings
Harvey AI

Join Harvey as a Senior Product Operations Manager to build and scale the evaluation engine for a global AI platform.

Harvey AI San Francisco Published 3 months ago
AI-first team
Town

Build eval systems and quality metrics for a personalized AI assistant at a startup in San Francisco.

Town San Francisco Published 1 week ago
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Distyl AI

Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.

Distyl AI San Francisco $150k–$250k/yr Published 2 months ago
70% coding
OpenRouter

Conduct original research on large language models to advance understanding and routing optimization at OpenRouter.

OpenRouter Remote (US) Published 1 month ago
Flexible on stack
Handshake

Join Handshake as an AI Model Policy Trainer to evaluate image generation models and shape AI's impact on careers.

Handshake Seattle, WA $45–$55/hr Published 1 week ago
Protege

Join Protege as a Forward Deployed Machine Learning Engineer to build the technical foundation for AI training data evaluations.

Protege Remote Published 1 month ago
Mind Robotics

Join Mind Robotics as a Research & Modeling Engineer to build and train core models for real-world robotic systems.

Mind Robotics Palo Alto Published 7 months ago
AI-first team