"evaluation systems" Jobs
1547 open tech roles matching “evaluation systems”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 1547 results
Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.
Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.
Build and improve the technical foundations for Answer Quality at Perplexity AI, collaborating with data scientists and engineers.
Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.
Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.
Join Anthropic as a Cyber Evaluations Engineer to design and run evaluations for AI systems, ensuring their safety and robustness.
Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.
Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.
Build eval systems and quality metrics for a personalized AI assistant at a startup in San Francisco.
Own end-to-end problems in building and improving AI agent learning infrastructure at Judgment Labs.
Lead the design and implementation of performance management programs at Reflection AI to support talent development and organizational growth.
Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.
Own the quality bar for human evaluations at a fast-scaling AI company transforming legal services.
Own and build internal products to support model evaluation and data management in a hybrid role at Twelve Labs.
Join Distyl AI as an Applied AI Researcher to redefine software usage and drive innovative benchmarking in AI systems.
Join Anthropic as a Product Designer to build and evaluate prompts for AI systems, ensuring alignment with user expectations and safety.