"evaluation systems" Jobs

528 open tech roles matching “evaluation systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, SQL. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 528 results

Cursor

Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.

Cursor San Francisco Published 3 months ago
Heavy meetings
Cartesia

Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.

Cartesia *HQ - San Francisco, CA Published 10 months ago
AI-first team
Reflection AI

Lead the design and implementation of performance management programs at Reflection AI to support talent development and organizational growth.

Reflection AI San Francisco, CA Published 1 month ago
Anthropic

Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.

Anthropic San Francisco, CA $500k–$850k/yr Published 1 month ago
Flexible on stack
Reliable Robotics Corporation

Lead and grow the Simulation and Test Systems team to enhance aviation safety through advanced simulation and testing technologies.

Reliable Robotics Corporation Mountain View, CA Published 1 week ago
Flexible on stack Heavy meetings
Handshake

Lead the Quality team at Handshake to enhance AI output reliability and evaluation accuracy in a fast-growing AI data business.

Handshake San Francisco, CA Published 2 months ago
Heavy meetings
Cursor

Lead a team of engineers to build infrastructure for training and evaluating ML models in a flat, innovative organization.

Cursor San Francisco Published 2 months ago
Heavy meetings
Benchling

Lead the design and development of full-stack systems for scientific workflows in a biotech AI platform.

Benchling San Francisco, CA Published 1 month ago
Flexible on stack
LaunchDarkly

Lead the AgentControl Evaluations team at LaunchDarkly, focusing on offline evaluations and AI-powered systems.

LaunchDarkly Remote - US $163k–$263.7k/yr Published 2 weeks ago
Flexible on stack Heavy meetings
baseten

Lead the finance systems strategy and implementation at a rapidly growing AI company.

baseten San Francisco Published 1 week ago
Discord

Lead a team of engineers to enhance automated review systems and improve safety processing at Discord.

Discord San Francisco Bay Area $248k–$279k/yr Published 1 month ago
AI-first team Heavy meetings
Varda Space

Lead a multidisciplinary team to ensure flight readiness of vehicles at Varda Space, a startup focused on commercial space infrastructure.

Varda Space El Segundo, California, United States $170k–$198k/yr Published 2 months ago
Flexible on stack Heavy meetings
Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 2 months ago
Applied Intuition

Lead the perception model team for autonomous vehicles at a rapidly growing AI infrastructure company.

Applied Intuition Sunnyvale Published 3 months ago
Flexible on stack
Benchling

Lead the Lab Notebook and Enterprise Lifecycle teams as a senior technical authority to drive impactful software solutions in biotech.

Benchling San Francisco, CA Published 1 month ago
Flexible on stack
Neko Health

Lead quality assurance for medical device hardware and firmware, ensuring compliance and safety in a fast-paced startup environment.

Neko Health Stockholm Published 5 months ago
Twelve Labs

Lead the implementation of business systems and analytics architecture at Twelve Labs, a growth-stage AI company.

Twelve Labs San Francisco Published 3 weeks ago
Astro Mechanica

Lead the development and integration of mission system architectures for advanced tactical autonomous aircraft at Astro Mechanica.

Astro Mechanica Remote Published 11 months ago
Perplexity AI

Lead the design systems at Perplexity AI, bridging design and engineering to create a coherent product experience across platforms.

Perplexity AI San Francisco Published 2 months ago
Flexible on stack
Datadog

Lead the Evaluation & Annotation team at Datadog, focusing on AI model evaluation and human annotation tooling.

Datadog Paris, France Published 3 months ago
Heavy meetings