Direct from source · No middlemen
5642 open positions · Updated 1 week ago
Showing 20 of 5642 positions
Search with filters →Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.
Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.
Join Triomics as an ML Evaluation Engineer to ensure model quality and stability in clinical AI systems.
Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.
Join Sarvam as a Data Scientist to design evaluation frameworks for AI outputs in high-stakes domains.
Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.
Join Anthropic as a Cyber Evaluations Engineer to design and run evaluations for AI systems, ensuring their safety and robustness.
Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.
Join Elastic as an AI QA & Evaluation Engineer to validate and test AI infrastructure and solutions in a fully remote environment.
Join Manex AI as a Line Impact Intern to contribute to customer projects in a high-growth AI startup environment.
Join Reflection AI as a hands-on technical staff member to enhance model performance through data-driven evaluations and feedback loops.
Join Reflection AI as a Research Program Manager to build foundational infrastructure for model evaluations and safety in AI.
Join Anthropic as a Product Designer to build and evaluate prompts for AI systems, ensuring alignment with user expectations and safety.
Join Reflection AI as a Research Software Engineer to build secure infrastructure for sensitive model evaluations in a fast-paced startup environment.
Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.
Lead the evaluation of Figma's AI-powered experiences to ensure quality and effectiveness in product features.
Build specialized evals to improve answer quality across Perplexity's products in a high-impact data science role.