Machine Learning Engineer — Model Evaluation & Experimentation
$60-$90 / hr
$60-$90

Location requirements
Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.
1. Overview
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced machine learning practitioners to act as ground-truth experts for model evaluation and experimentation. You will author complex, multi-step ML tasks — for example, taking a vague research idea like "modify how an RL reward is computed," implementing the change, running the training experiment, and analyzing the results to determine success — and verify exactly where frontier models fall short.
Each task represents one to two days of continuous, focused effort and spans multiple technical skills: implementation, experiment setup and execution, and rigorous analysis. You will work in a tight feedback loop with the lab's researchers.
This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.
2. Key Responsibilities
-
Design tasks: Turn real ML research ideas — like tweaking how an RL reward is computed — into well-defined, multi-step tasks.
-
Run experiments: Implement changes, run training experiments, and analyze the results to show what a correct solution looks like.
-
Explore RL ideas: Build some of your tasks around reinforcement-learning basics such as reward functions and training behavior.
-
Evaluate models: See how frontier models handle your tasks, and note where and why they fall short.
-
Work as a team: Compare notes with researchers and fellow experts so tasks stay consistent, rigorous, and fair.
3. Core Qualifications
-
MSc or PhD in machine learning, computer science, or another STEM field, or equivalent practical experience in a research-heavy domain.
-
1+ years of experience in a research or research-engineering role.
-
Hands-on experience training and evaluating ML models and running experiments end-to-end — experiment setup, execution, and analysis — not just using ML libraries superficially.
-
Strong familiarity with large language models: their capabilities, limitations, and evaluation techniques.
-
Working proficiency in Python and Git, with comfort in both scripting and notebook environments.
-
Basic understanding of reinforcement learning (reward functions, policy training) is preferred.
-
Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.
-
A perfectionist mindset: high attention to detail, creativity in task design, strong written communication, and the ability to work independently through ambiguous, open-ended problems.
-
Ability to engage reliably for approximately 35 hours per week.
About Cincinnatus LLC
Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.
Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows.
Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.
Equal Employment Opportunity
Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.
Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.
We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.
Earn up to $1,440 by referring
Posted 8 hours ago