"evaluators" Jobs

1875 open tech roles matching “evaluators”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 1875 results

Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 2 months ago
Nuna

Own the evaluation system for Nuna's AI health coach, building testing harnesses and infrastructure to ensure safety and quality.

Nuna San Francisco Published 1 month ago
Figma

Lead the evaluation of Figma's AI-powered experiences to ensure quality and effectiveness in product features.

Figma San Francisco, CA • New York, NY • United States $258k–$348k/yr Published 2 months ago
Vals AI

Join Vals AI as a researcher to design and build next-gen AI benchmarks in a fast-paced, innovative environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Vals AI

Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Fieldguide

Join Fieldguide as a Senior AI Engineer to build reliable AI agents for audit workflows in a remote-first environment.

Fieldguide San Francisco, CA or Remote (USA) Published 4 months ago
Flexible on stack
Snorkel AI

Lead a team of researchers to enhance data evaluation and analysis methods for AI model performance at Snorkel AI.

Snorkel AI New York City, NY (Hybrid); San Francisco, CA (Hybrid); United States (Remote) $275k–$425k/yr Published 3 months ago
AI-first team
Handshake

Lead the Quality team at Handshake to enhance AI output reliability and evaluation accuracy in a fast-growing AI data business.

Handshake San Francisco, CA Published 2 months ago
Heavy meetings
Reflection AI

Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.

Reflection AI San Francisco, CA Published 8 months ago
Flexible on stack
Judgment Labs

Own end-to-end problems in building and improving AI agent learning infrastructure at Judgment Labs.

Judgment Labs San Francisco Published 2 months ago
Cartesia

Join Cartesia as a Product and Research Operations Manager to design and scale a global evaluation workforce for AI.

Cartesia *HQ - San Francisco, CA Published 2 weeks ago
AI-first team
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Judgment Labs

Join Judgment Labs as a Senior Backend Engineer to build infrastructure for AI agents in a fast-paced, onsite environment in San Francisco.

Judgment Labs San Francisco Published 3 months ago
Sesame

Join Sesame as a Research Engineer to innovate in NLP, Speech, and Computer Vision with a focus on deep learning.

Sesame San Francisco Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Research Engineer to build evaluation instruments for AI development in a remote-friendly environment.

Anthropic Remote-Friendly (Travel Required) | San Francisco, CA $350k–$850k/yr Published 5 days ago
Flexible on stack
Ambience Healthcare

Build and improve AI systems for clinical products as a Senior Machine Learning Engineer in a hybrid role at Ambience Healthcare.

Ambience Healthcare San Francisco $225k–$300k/yr Published 1 month ago
Flexible on stack 70% coding
Twelve Labs

Own and build internal products to support model evaluation and data management in a hybrid role at Twelve Labs.

Twelve Labs San Francisco Published 1 month ago
Flexible on stack
Anthropic

Drive model launches and improve coding performance as a Product Manager on Claude Code's model performance team.

Anthropic San Francisco, CA | Seattle, WA $305k–$460k/yr Published 3 months ago