"model evaluation" Jobs
4280 open tech roles matching “model evaluation”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, SQL. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 4280 results
Join Anthropic as a Staff+ Software Engineer to build and maintain the RL Data Platform for reliable AI systems.
Design and maintain backend services and AI systems while collaborating with research and product teams in a fast-paced startup environment.
Join a senior team to develop high-fidelity material characterization methods leveraging AI and automation in a groundbreaking lab.
Build autonomous AI agents for go-to-market operations in a senior engineering role at Anthropic.
Join Together AI as a Platform Engineer to build and improve infrastructure for model customization and evaluation in a collaborative environment.
Join Neko Health as an ML Ops Engineer to lead the ML infrastructure for preventive care and early detection.
Join HoneyBook as a Senior Machine Learning Engineer to build and maintain AI-powered systems in a fast-paced environment.
Lead the technical strategy for Ads ML engineer lifecycle at Reddit, enhancing ML development processes and mentoring engineers.
Join OpenRouter as a Forward Deployed Engineer to help customers implement and scale AI solutions effectively.
Build and own the infrastructure and pipelines for machine learning models in production at a fast-growing startup.
Join Handshake as an AI Model Policy Trainer to evaluate AI responses in mental health conversations.
Join Reddit as a Machine Learning Engineer to optimize ads auction and bidding systems for a high-impact revenue engine.
Join RobCo as a Senior Solution Development Engineer to design and validate robotic systems for autonomous industrial solutions.
Lead the unification of large datasets to power generative audio models in a fast-moving team at Udio.
Join Anthropic as a Staff Researcher to shape AI capabilities for cybersecurity, focusing on practical applications for non-experts.
Lead a data team at Cartesia to enhance the quality of multimodal AI training data and infrastructure.