"evaluation frameworks" Jobs

2731 open tech roles matching “evaluation frameworks”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, AWS. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 2731 results

Handshake

Join Handshake as an AI Model Policy Trainer to evaluate image generation models and shape AI's impact on careers.

Handshake Seattle, WA $45–$55/hr Published 1 month ago
Abridge

Join Abridge as a Research Scientist to evaluate the impact of ambient AI on healthcare outcomes in a fast-paced startup environment.

Abridge NYC Office Published 7 months ago
Arena

Own the full lifecycle of revenue generation in a senior role at Arena, focusing on AI model evaluation.

Arena SF Bay Area Published 1 month ago
Reflection AI

Join Reflection AI as a Research Program Manager to build foundational infrastructure for model evaluations and safety in AI.

Reflection AI New York, NY Published 5 months ago
Reflection AI

Join Reflection AI as a Research Software Engineer to build secure infrastructure for sensitive model evaluations in a fast-paced startup environment.

Reflection AI New York, NY Published 2 months ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY | Washington, DC $230k–$270k/yr Published 7 months ago
Fieldguide

Lead the integration of audit methodologies in a remote-first startup focused on enhancing trust in global commerce.

Fieldguide San Francisco, CA or Remote (USA) Published 1 month ago
Rox

Join Rox as a Founding Applied Research Engineer to shape the future of applied AI with a focus on real-world production challenges.

Rox San Francisco Published 4 months ago
Arena

Join Arena as a Member of Technical Staff to build data infrastructure for evaluating AI model performance.

Arena SF Bay Area Published 9 months ago
Flexible on stack
Hilbert

Join Hilbert as an AI Engineer to build production-grade AI systems that drive real enterprise outcomes.

Hilbert Türkiye Published 3 weeks ago
Flexible on stack 70% coding
Fieldguide

Lead the financial audit strategy for EMEA at Fieldguide, driving platform localization and partnerships.

Fieldguide Remote (UK) Published 1 week ago
OpenEvidence

Lead clinical research and evaluation for OpenEvidence's AI platform in global health, focusing on low-resource settings.

OpenEvidence Miami Published 2 weeks ago
Reflection AI

Serve as a Subject Matter Expert on election integrity and fraud, ensuring Reflection's models meet safety and compliance standards.

Reflection AI New York, NY Published 2 months ago
Cognition

Join Cognition as a QA Engineer to ensure product quality across AI workflows and contribute to innovative AI solutions.

Cognition India Published 7 months ago
Flexible on stack AI-first team
Fireworks AI

Join Fireworks AI as a Member of Technical Staff to enhance model evaluation and fine-tuning workflows in a fast-paced generative AI environment.

Fireworks AI San Mateo Published 11 months ago
Hilbert

Build production-grade AI systems end-to-end at Hilbert, a fast-growing company in San Francisco.

Hilbert San Francisco Published 3 weeks ago
Flexible on stack 70% coding
Fieldguide

Join Fieldguide as a Senior AI Engineer to build reliable AI agents for audit workflows in a remote-first environment.

Fieldguide San Francisco, CA or Remote (USA) Published 5 months ago
Flexible on stack
Edra

Join Edra as an AI Engineer to build complex LLM-based systems that enhance enterprise AI processes.

Edra London Published 11 months ago
Arena

Join Arena as an Associate General Counsel to lead privacy compliance and advise on AI product development in a mission-driven startup.

Arena SF Bay Area Published 2 months ago
AI-first team
Crucibl

Join Crucibl as a Member of Technical Staff to push the frontier of judgment in AI and shape the future of decision-making.

Crucibl San Francisco, CA Published 2 months ago