SWE-Bench Task Auditor
$70-$90 / hr
$70-$90

Location requirements
Evaluate the quality, correctness, and reproducibility of software-engineering benchmark tasks used to train and evaluate a frontier AI lab's models. You'll assess repository-level tasks, reference patches, test harnesses, and grading integrity — and provide clear, rubric-based written feedback.
Basic Qualifications • 3+ years professional software engineering • Real open-source contribution or maintainer experience (merged PRs, committer / maintainer roles) • Strong ability to audit reference patches, test runners, and Docker isolation, and to detect answer leakage / reward hacking • Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C++)
Preferred Qualifications • Familiarity with SWE-Bench (Verified) or similar repository benchmarks • Maintainer history on major Python OSS (Django, Flask, scikit-learn, sympy, pytest, etc.) • Prior code-review or task-grading experience
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Contract and Payment Terms
- You will be engaged as an independent contractor.
- This is a fully remote role that can be completed on your own schedule.
- Projects can be extended, shortened, or concluded early depending on needs and performance.
- Your work at Mercor will not involve access to confidential or proprietary information from any employer, client, or institution.
- Payments are weekly on Stripe or Wise based on services rendered.
- Please note: We are unable to support H1-B or STEM OPT candidates at this time.
About Mercor
Mercor partners with leading AI labs and enterprises to train frontier models using human expertise. You will work on projects that focus on training and enhancing AI systems. You will be paid competitively, collaborate with leading researchers, and help shape the next generation of AI systems in your area of expertise.
Earn up to $1,440 by referring
Posted 3 days ago