"ai model evaluation" Jobs

1266 open tech roles matching “ai model evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 1266 results

Mirendil

Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
AI-first team
Cartesia

Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.

Cartesia *HQ - San Francisco, CA Published 10 months ago
AI-first team
Vals AI

Join Vals AI as an Evaluations Engineer to evaluate LLM models and contribute to industry-leading benchmarks.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Vals AI

Join Vals AI as a researcher to design and build next-gen AI benchmarks in a fast-paced, innovative environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Anthropic
Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY $320k–$485k/yr Published 4 months ago
Town

Build eval systems and quality metrics for a personalized AI assistant at a startup in San Francisco.

Town San Francisco Published 2 weeks ago
Reflection AI

Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.

Reflection AI San Francisco, CA Published 9 months ago
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Exa

Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.

Exa San Francisco, California Published 11 months ago
Flexible on stack
Anthropic

Join Anthropic as a Cyber Evaluations Engineer to design and run evaluations for AI systems, ensuring their safety and robustness.

Anthropic San Francisco, CA | Washington, DC $300k–$405k/yr Published 1 week ago
Flexible on stack
Distyl AI

Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.

Distyl AI San Francisco $150k–$250k/yr Published 2 months ago
70% coding
krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Harvey AI

Join Harvey as a Senior Product Operations Manager to build and scale the evaluation engine for a global AI platform.

Harvey AI San Francisco Published 3 months ago
AI-first team
Anthropic

Drive model launches and improve coding performance as a Product Manager on Claude Code's model performance team.

Anthropic San Francisco, CA | Seattle, WA $305k–$460k/yr Published 3 months ago
Cursor

Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.

Cursor San Francisco Published 3 months ago
Heavy meetings
Cartesia

Join Cartesia as a Researcher to advance neural network architecture design in a collaborative, innovative environment.

Cartesia *HQ - San Francisco, CA Published 1 year ago
Flexible on stack
Benchling

Join Benchling as a Research Engineer to improve AI models for scientific applications in a fast-paced, collaborative environment.

Benchling San Francisco, CA Published 3 weeks ago
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in large language models within a fast-paced startup.

Preference Model San Francisco, United States Published 2 days ago
Flexible on stack
Anthropic

Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.

Anthropic San Francisco, CA $500k–$850k/yr Published 1 month ago
Flexible on stack