"evaluation pipelines" Jobs
2739 open tech roles matching “evaluation pipelines”, taken straight from company career pages — not reposted from other job boards. Most in demand right now: Python, AI/ML, AWS. Every listing is re-checked daily and closed roles are removed.
Showing 20 of 2739 results
Join Protege as a Forward Deployed Machine Learning Engineer to build the technical foundation for AI training data evaluations.
Build an AI-native regulatory intelligence system in a 12-week residency focused on operational decisions in regulated environments.
Join Reflection AI as a Data Quality Engineer to ensure high-quality data for model training in a mission-driven research lab.
Join Arena as a Member of Technical Staff to build data infrastructure for evaluating AI model performance.
Join MaintainX as a Senior SDET to build shared testing infrastructure and AI-assisted tooling for high-quality software delivery.
Own quality for an AI-enabled decision-support system as a QA Engineer in a fully remote role.
Join MaintainX as a Senior SDET to build a shared quality platform and enhance AI-powered testing across engineering teams.
Join Reflection AI as a Forward Deployed Engineer to fine-tune models and work directly with enterprise customers in a dynamic startup environment.
Join EliseAI as a Senior DevOps Engineer to build and maintain infrastructure that supports reliable software deployment in a fast-paced environment.
Join Mirendil as a research engineer to build evaluation infrastructure for frontier AI models.
Lead quality management for AI deliverables at Lilt, ensuring high standards in a hybrid work environment.
Join Anthropic as a Tech Lead to build reliable AI evaluation systems in a hybrid work environment.
Join Iambic Therapeutics as a Software Engineer to develop agentic data pipelines for biomedical data using LLMs.
Join Cantina as a Machine Learning Engineer to lead audio model evaluations and shape the future of social AI.
Join Judgment Labs as an Applied AI Engineer to build self-improving AI systems using real-world agent interaction data.
Join Cartesia as a lead researcher to design evaluation frameworks for next-generation AI models.
Lead a team of engineers to build infrastructure for training and evaluating ML models in a flat, innovative organization.
Build a scalable Experiment Platform at Poolside AI, working with a remote-first team to enhance research workflows.
Join Postman as an AI Engineer Intern to work on large-scale AI systems with mentorship from senior engineers.