"evaluation infrastructure" Jobs
866 open tech roles matching “evaluation infrastructure”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: AI/ML, Python, TypeScript. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 866 results
Lead the Evals team at Cursor to create high-signal evaluation datasets and tools for coding agents.
Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.
Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.
Join Judgment Labs as a Senior Backend Engineer to build infrastructure for AI agents in a fast-paced, onsite environment in San Francisco.
Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.
Join MaintainX as a Senior SDET to build a shared quality platform and enhance AI-powered testing across engineering teams.
Build and improve the technical foundations for Answer Quality at Perplexity AI, collaborating with data scientists and engineers.
Join Vals AI as a mid-level engineer to build and maintain a platform for evaluating LLMs at scale in a dynamic startup environment.
Lead a team of engineers to build infrastructure for training and evaluating ML models in a flat, innovative organization.
Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.
Build production-grade AI systems end-to-end at Hilbert, a fast-growing company in San Francisco.
Lead the product strategy for Abridge's AI/ML evaluation platform, ensuring quality and efficiency across multiple product teams.
Own end-to-end problems in building and improving AI agent learning infrastructure at Judgment Labs.
Join Baseten as a Software Engineer to drive model performance systems at the intersection of HPC and LLM engineering.
Join Preference Model as a Senior ML Infrastructure Engineer to build scalable infrastructure for post-training research on large language models.
Lead the Quality team at Handshake to enhance AI output reliability and evaluation accuracy in a fast-growing AI data business.