"evaluation framework" Jobs

2672 open tech roles matching “evaluation framework”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, AWS. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 2672 results

Handshake

Join Handshake as a Member of Technical Staff to define and build evaluation frameworks for frontier AI systems.

Handshake San Francisco, CA Published 1 week ago
Flexible on stack
Exa

Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.

Exa San Francisco, California Published 11 months ago
Flexible on stack
Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 3 months ago
Sarvam AI

Join Sarvam as a Data Scientist to design evaluation frameworks for AI outputs in high-stakes domains.

Sarvam AI Bengaluru Published 5 months ago
Flexible on stack
Cartesia

Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.

Cartesia *HQ - San Francisco, CA Published 11 months ago
AI-first team
DEFCON AI

Join DEFCON AI as a Model Test and Measurement Engineer to ensure AI systems perform reliably and transparently in a fully remote role.

DEFCON AI Remote, USA $150k–$190k/yr Published 6 days ago
Arena

Join Arena as a Software Engineer to build core infrastructure for AI model evaluation in a fast-paced startup environment.

Arena SF Bay Area Published 2 weeks ago
70% coding
Anthropic

Join Anthropic as a Cyber Evaluations Engineer to design and run evaluations for AI systems, ensuring their safety and robustness.

Anthropic Remote-Friendly, United States; San Francisco, CA | Washington, DC $300k–$405k/yr Published 1 month ago
Flexible on stack
Triomics

Join Triomics as an ML Evaluation Engineer to ensure model quality and stability in clinical AI systems.

Triomics India Office Published 3 months ago
Flexible on stack
Vals AI

Join Vals AI as a researcher to design and build next-gen AI benchmarks in a fast-paced, innovative environment.

Vals AI San Francisco Published 3 months ago
Flexible on stack
Arena

Join Arena as a Machine Learning Scientist to evaluate AI models and contribute to impactful research in a collaborative environment.

Arena SF Bay Area Published 9 months ago
Flexible on stack 70% coding
Mirendil

Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.

Mirendil San Francisco $300k–$400k/yr Published 3 months ago
AI-first team
Gong

Join Gong as a Senior Backend Engineer to build foundational AI frameworks that empower revenue teams.

Gong Tel Aviv Published 3 months ago
Flexible on stack
Figma

Lead the evaluation of Figma's AI-powered experiences to ensure quality and effectiveness in product features.

Figma San Francisco, CA • New York, NY • United States $258k–$348k/yr Published 2 months ago
Reflection AI

Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.

Reflection AI San Francisco, CA Published 9 months ago
Town

Build eval systems and quality metrics for a personalized AI assistant at a startup in San Francisco.

Town San Francisco Published 1 month ago
Arena

Join Arena as an Applied AI Engineer to integrate AI models and build solutions for cutting-edge AI teams.

Arena SF Bay Area Published 3 weeks ago
Flexible on stack
Arena

Join Arena as a Product Solutions Engineer to work closely with leading AI labs and build impactful solutions.

Arena SF Bay Area Published 7 months ago
Flexible on stack 70% coding
Reflection AI

Join Reflection AI as a Research Program Manager to build foundational infrastructure for model evaluations and safety in AI.

Reflection AI New York, NY Published 5 months ago
Distyl AI

Join Distyl AI as an Applied AI Researcher to redefine software usage and drive innovative benchmarking in AI systems.

Distyl AI San Francisco $150k–$250k/yr Published 11 months ago