"evaluation infrastructure" Jobs

866 open tech roles matching “evaluation infrastructure”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 866 results

Cursor

Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.

Cursor San Francisco Published 3 months ago
Heavy meetings
Cartesia

Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.

Cartesia *HQ - San Francisco, CA Published 10 months ago
AI-first team
Anthropic

Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.

Anthropic San Francisco, CA $500k–$850k/yr Published 1 month ago
Flexible on stack
Vals AI

Join Vals AI as a researcher to design and build next-gen AI benchmarks in a fast-paced, innovative environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Judgment Labs

Join Judgment Labs as a Senior Backend Engineer to build infrastructure for AI agents in a fast-paced, onsite environment in San Francisco.

Judgment Labs San Francisco Published 3 months ago
Mirendil

Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
AI-first team
MaintainX

Join MaintainX as a Senior SDET to build a shared quality platform and enhance AI-powered testing across engineering teams.

MaintainX San Francisco Published 6 days ago
Flexible on stack
Perplexity AI

Build and improve the technical foundations for Answer Quality at Perplexity AI, collaborating with data scientists and engineers.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Vals AI

Join Vals AI as an Evaluations Engineer to evaluate LLM models and contribute to industry-leading benchmarks.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Vals AI

Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Cursor

Lead a team of engineers to build infrastructure for training and evaluating ML models in a flat, innovative organization.

Cursor San Francisco Published 2 months ago
Heavy meetings
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Hilbert

Build production-grade AI systems end-to-end at Hilbert, a fast-growing company in San Francisco.

Hilbert San Francisco, United States Published 15 hours ago
Flexible on stack 70% coding
Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 2 months ago
Anthropic
Anthropic Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY $320k–$485k/yr Published 4 months ago
Judgment Labs

Own end-to-end problems in building and improving AI agent learning infrastructure at Judgment Labs.

Judgment Labs San Francisco Published 2 months ago
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 3 days ago
Flexible on stack
Handshake

Lead the Quality team at Handshake to enhance AI output reliability and evaluation accuracy in a fast-growing AI data business.

Handshake San Francisco, CA Published 2 months ago
Heavy meetings