Adverise with usGet your job listing or Product in front of thousands of AI trainers.
Contact us
Role Overview
Mercor is running a short, structured evaluation pilot to measure how frontier AI models perform on realistic professional work — and how experienced practitioners judge that work. We are hiring 25 experts across four business domains. Each expert completes one task, then evaluates five model attempts at that same task. This is a one-time study, not ongoing production work.
What You'll Do
Time Commitment
Up to 13 hours total: 3–10 hours to solve the task, approximately 2 hours to rank and score the five outputs, and approximately 1 hour to provide feedback.
Study Conditions
Domains we are focused on
Who We're Looking For
Eligibility Restriction
You are not eligible for this pilot if you have worked on Project Alchemy in any capacity — task author, reviewer, world expert, or anyone who has had access to its source world data. Prior exposure to that material would invalidate the study.
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Mercor partners with leading AI labs and enterprises to train frontier models using human expertise. You will work on projects that focus on training and enhancing AI systems. You will be paid competitively, collaborate with leading researchers, and help shape the next generation of AI systems in your area of expertise.