"llm evaluation" Jobs
496 open tech roles matching “llm evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 496 results
Join the Model Behavior team at Perplexity AI to shape AI product responses through context and prompt engineering.
Build and improve AI systems for clinical products as a Senior Machine Learning Engineer in a hybrid role at Ambience Healthcare.
Join the Model Behavior team to shape Perplexity’s AI products through context and prompt engineering.
Lead and mentor a team of Forward Deployed Engineers to optimize LLM inference workloads for Baseten customers.
Join Reflection AI as a Data Quality Engineer to ensure high data standards for AI model training and evaluation.
Join Crucibl as a Member of Technical Staff to push the frontier of judgment in AI and shape the future of decision-making.
Own the quality bar for human evaluations at a fast-scaling AI company transforming legal services.
Join Braintrust as a backend engineer to build infrastructure for cutting-edge AI development tools in a fast-paced environment.
Join Distyl AI as a Senior AI Engineer to design evaluation frameworks that enhance AI systems in production.
Join Krea as an ML Researcher to finetune diffusion models and enhance AI creative tools in a collaborative environment.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.
Join Langfuse as a Developer Relations Engineer to enhance community engagement and represent the company at key events across Europe and the US.
Lead a multidisciplinary research team to advance large-scale machine learning efficiency at Databricks.