"evaluation frameworks" Jobs
2737 open tech roles matching “evaluation frameworks”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, AWS. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 2737 results
Join Handshake as a Member of Technical Staff to define and build evaluation frameworks for frontier AI systems.
Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.
Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.
Join Arena as a Machine Learning Scientist to evaluate AI models and contribute to impactful research in a collaborative environment.
Join Sarvam as a Data Scientist to design evaluation frameworks for AI outputs in high-stakes domains.
Join Anthropic as a Cyber Evaluations Engineer to design and run evaluations for AI systems, ensuring their safety and robustness.
Join Arena as a Software Engineer to build core infrastructure for AI model evaluation in a fast-paced startup environment.
Join DEFCON AI as a Model Test and Measurement Engineer to ensure AI systems perform reliably and transparently in a fully remote role.
Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.
Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.
Join Triomics as an ML Evaluation Engineer to ensure model quality and stability in clinical AI systems.
Build eval systems and quality metrics for a personalized AI assistant at a startup in San Francisco.
Join Arena as a Product Solutions Engineer to work closely with leading AI labs and build impactful solutions.
Join Arena as an Applied AI Engineer to integrate AI models and build solutions for cutting-edge AI teams.
Join Distyl AI as an Applied AI Researcher to redefine software usage and drive innovative benchmarking in AI systems.
Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.
Lead the evaluation of Figma's AI-powered experiences to ensure quality and effectiveness in product features.