"evaluation pipelines" Jobs

790 open tech roles matching “evaluation pipelines”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 790 results

Exa

Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.

Exa San Francisco, California Published 11 months ago
Flexible on stack
Distyl AI

Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.

Distyl AI San Francisco $150k–$250k/yr Published 2 months ago
70% coding
Anthropic

Join Anthropic as a Product Designer to build and evaluate prompts for AI systems, ensuring alignment with user expectations and safety.

Anthropic San Francisco, CA $305k–$385k/yr Published 1 week ago
Flexible on stack
Perplexity AI

Build specialized evals to improve answer quality across Perplexity's products in a high-impact data science role.

Perplexity AI San Francisco Published 2 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to build continuous deployment infrastructure for AI products in a fast-growing team.

baseten San Francisco Published 1 month ago
Flexible on stack
Judgment Labs

Join Judgment Labs as a Senior Backend Engineer to build infrastructure for AI agents in a fast-paced, onsite environment in San Francisco.

Judgment Labs San Francisco Published 3 months ago
Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 2 months ago
Vals AI

Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Perplexity AI

Build and improve the technical foundations for Answer Quality at Perplexity AI, collaborating with data scientists and engineers.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Reflection AI

Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.

Reflection AI San Francisco, CA Published 8 months ago
Flexible on stack
baseten

Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.

baseten San Francisco Published 8 months ago
Flexible on stack
Reflection AI

Join Reflection AI as a Data Quality Engineer to ensure high-quality data for model training in a mission-driven research lab.

Reflection AI San Francisco, CA Published 8 months ago
Flexible on stack
MaintainX

Join MaintainX as a Senior SDET to build a shared quality platform and enhance AI-powered testing across engineering teams.

MaintainX San Francisco Published 6 days ago
Flexible on stack
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Mirendil

Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.

Mirendil San Francisco $300k–$400k/yr Published 2 months ago
AI-first team
Anthropic

Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.

Anthropic San Francisco, CA $500k–$850k/yr Published 1 month ago
Flexible on stack
Judgment Labs

Join Judgment Labs as an Applied AI Engineer to build self-improving AI systems using real-world agent interaction data.

Judgment Labs San Francisco Published 8 months ago
Cartesia

Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.

Cartesia *HQ - San Francisco, CA Published 10 months ago
AI-first team