"ai model evaluation" Jobs

1266 open tech roles matching “ai model evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 1266 results

Macroscope

Join Macroscope as an Applied ML Engineer to enhance machine learning systems in a collaborative startup environment.

Macroscope San Francisco $170k–$280k/yr Published 1 month ago
Flexible on stack
Preference Model

Join Preference Model as a Research Engineer to advance self-directed learning in machine learning with a focus on RL environments.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Vals AI

Lead the research efforts at Vals AI to advance the science of evaluation in the AI economy.

Vals AI San Francisco, United States Published 2 months ago
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Vals AI

Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Anthropic
Anthropic San Francisco, CA | New York City, NY | Washington, DC $230k–$270k/yr Published 6 months ago
Perplexity AI

Build specialized evals to improve answer quality across Perplexity's products in a high-impact data science role.

Perplexity AI San Francisco Published 2 months ago
Flexible on stack
Anthropic

Join Anthropic as a Product Designer to build and evaluate prompts for AI systems, ensuring alignment with user expectations and safety.

Anthropic San Francisco, CA $305k–$385k/yr Published 1 week ago
Flexible on stack
Cartesia

Join Cartesia as a Product and Research Operations Manager to design and scale a global evaluation workforce for AI.

Cartesia *HQ - San Francisco, CA Published 2 weeks ago
AI-first team
Snorkel AI

Lead a team of researchers to enhance data evaluation and analysis methods for AI model performance at Snorkel AI.

Snorkel AI New York City, NY (Hybrid); San Francisco, CA (Hybrid); United States (Remote) $275k–$425k/yr Published 3 months ago
AI-first team
Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 2 months ago
Harvey

Lead the design and development of systems powering AI requests at Harvey, ensuring high reliability and operational excellence.

Harvey San Francisco $193.4k–$290k/yr Published 1 week ago
Flexible on stack 60% coding
Cartesia

Join Cartesia as a Researcher to enhance multimodal models through innovative post-training methods and alignment techniques.

Cartesia *HQ - San Francisco, CA Published 10 months ago
Cartesia

Join Cartesia as an Applied Researcher to enhance generative audio models by bridging customer needs with innovative research.

Cartesia *HQ - San Francisco, CA Published 1 month ago
Reflection AI

Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.

Reflection AI San Francisco, CA Published 8 months ago
Flexible on stack
Preference Model

Join Preference Model as a Machine Learning Engineer to develop low-level reinforcement learning environments in a fast-paced startup.

Preference Model San Francisco Published 2 weeks ago
Flexible on stack
Perplexity AI

Join the Model Behavior team at Perplexity AI to shape AI product responses through context and prompt engineering.

Perplexity AI San Francisco Published 1 month ago
Twelve Labs

Own and build internal products to support model evaluation and data management in a hybrid role at Twelve Labs.

Twelve Labs San Francisco Published 1 month ago
Flexible on stack