"llm evaluation" Jobs

1533 open tech roles matching “llm evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, TypeScript. Every listing is re-checked daily and closed roles are removed.

Showing 20 of 1533 results

krea.ai

Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.

krea.ai San Francisco Published 1 week ago
Flexible on stack
Lilt

Join LILT as a Forward Deployed Engineer to integrate AI solutions for complex clients and enhance global communication.

Lilt London, UK Published 1 month ago
Flexible on stack
Crosby

Join Crosby as an AI Developer Experience Engineer to enhance developer productivity in building AI features for legal systems.

Crosby New York City Published 2 months ago
Flexible on stack
Langfuse

Join Langfuse as a Senior Product Engineer to build impactful open-source AI applications in a small, dynamic team.

Langfuse Europe Published 6 months ago
Flexible on stack
LaunchDarkly

Lead the AgentControl Evaluations team at LaunchDarkly, focusing on offline evaluations and AI-powered systems.

LaunchDarkly Remote - US $163k–$263.7k/yr Published 2 weeks ago
Flexible on stack Heavy meetings
Edra

Join Edra as a Forward Deployed AI Engineer to build LLM-powered systems for enterprise customers in a hands-on engineering role.

Edra London Published 1 month ago
70% coding
Handshake

Stress-test large language models by crafting adversarial prompts to expose vulnerabilities in AI systems.

Handshake Seattle, WA Published 3 days ago
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco, United States Published 1 day ago
Flexible on stack
Preference Model

Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.

Preference Model San Francisco Published 2 days ago
Flexible on stack
Sphere

Lead the development of TRAM, an AI reasoning model for interpreting global trade law in a fast-paced, onsite environment.

Sphere San Francisco HQ Published 11 months ago
LangChain

Lead the engineering team building LangSmith, an observability and evaluation platform for LLM applications.

LangChain Boston, MA $200k–$240k/yr Published 2 days ago
Flexible on stack Heavy meetings
Abridge

Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.

Abridge SF Office Published 2 months ago
Protege

Join Protege as a Forward Deployed Machine Learning Engineer to build the technical foundation for AI training data evaluations.

Protege Remote Published 1 month ago
Langfuse

Join Langfuse as a Senior Software Engineer to shape the developer experience of our widely-used SDKs in the AI space.

Langfuse Europe Published 2 weeks ago
Flexible on stack