"llm evaluation" Jobs

496 open tech roles matching “llm evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 496 results

Perplexity AI

Build specialized evals to improve answer quality across Perplexity's products in a high-impact data science role.

Perplexity AI San Francisco Published 2 months ago
Flexible on stack
Vals AI

Join Vals AI as an Evaluations Engineer to evaluate LLM models and contribute to industry-leading benchmarks.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Vals AI

Join Vals AI as a researcher to design and build next-gen AI benchmarks in a fast-paced, innovative environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Vals AI

Lead the research efforts at Vals AI to advance the science of evaluation in the AI economy.

Vals AI San Francisco, United States Published 2 months ago
Reflection AI

Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.

Reflection AI San Francisco, CA Published 4 months ago
Exa

Join Exa as an ML evals engineer to design and build evaluation frameworks for a groundbreaking AI search engine.

Exa San Francisco, California Published 11 months ago
Flexible on stack
Perplexity AI

Build and improve the technical foundations for Answer Quality at Perplexity AI, collaborating with data scientists and engineers.

Perplexity AI San Francisco Published 5 months ago
Flexible on stack
Goodfire

Join Goodfire as an Account Executive to drive strategic partnerships in the LLM market with a focus on AI safety and capability.

Goodfire San Francisco, CA Published 2 months ago
MaintainX

Own the LLMX roadmap as a Staff Product Manager at MaintainX, driving AI platform adoption and quality in a hybrid role.

MaintainX San Francisco Published 3 months ago
MaintainX

Own the LLMX roadmap as a Staff Product Manager at MaintainX, driving AI platform adoption and quality.

MaintainX Toronto Published 1 month ago
Perplexity

Join Perplexity as a Senior ML Engineer to design and optimize recommendation systems for personalized user experiences.

Perplexity San Francisco Published 3 months ago
70% coding
Vals AI

Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.

Vals AI San Francisco, United States Published 2 months ago
Flexible on stack
Reflection AI

Conduct critical analysis and develop evaluation frameworks to improve AI model capabilities in a fast-paced startup environment.

Reflection AI San Francisco, CA Published 9 months ago
Perplexity AI

Join Perplexity AI as a senior ML Engineer to design and optimize recommendation systems for personalized user experiences.

Perplexity AI San Francisco Published 3 months ago
Lightfield

Join Lightfield as a Staff Software Engineer to leverage LLMs in building innovative AI-powered CRM solutions.

Lightfield HQ: San Francisco Published 4 months ago
Flexible on stack 70% coding
Datadog
Datadog Boston, Massachusetts, USA; Denver, Colorado, USA; New York, New York, USA; San Francisco, California, USA $151.5k–$222k/yr Published 6 months ago
Figma

Lead the evaluation of Figma's AI-powered experiences to ensure quality and effectiveness in product features.

Figma San Francisco, CA • New York, NY • United States $258k–$348k/yr Published 2 months ago