"evaluation pipelines" Jobs
790 open tech roles matching “evaluation pipelines”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 790 results
Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.
Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.
Join Anthropic as a Product Designer to build and evaluate prompts for AI systems, ensuring alignment with user expectations and safety.
Build specialized evals to improve answer quality across Perplexity's products in a high-impact data science role.
Join Baseten as a Software Engineer to build continuous deployment infrastructure for AI products in a fast-growing team.
Join Judgment Labs as a Senior Backend Engineer to build infrastructure for AI agents in a fast-paced, onsite environment in San Francisco.
Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.
Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.
Build and improve the technical foundations for Answer Quality at Perplexity AI, collaborating with data scientists and engineers.
Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.
Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.
Join Reflection AI as a Data Quality Engineer to ensure high-quality data for model training in a mission-driven research lab.
Join MaintainX as a Senior SDET to build a shared quality platform and enhance AI-powered testing across engineering teams.
Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.
Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.
Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.
Join Judgment Labs as an Applied AI Engineer to build self-improving AI systems using real-world agent interaction data.
Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.