LLM Red Team Specialist — Failure Modes & Edge Cases

$60-$90 / hr

Full-time position
United States

$60-$90

per hour

Mercor logo
Posted by Mercor

Location requirements

USA

Join a leading AI lab's cutting-edge GenAI team to be at the core of the AI revolution, where your expertise fuels the development of the most advanced AI models.

1. Overview

A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs specialists who can find where those models break. Working in a red-teaming setup, you will design and probe complex, multi-step tasks to expose vulnerabilities, edge cases, and failure modes in frontier AI systems — the places where a model looks competent but is quietly wrong.

Each task represents one to two days of continuous, focused effort and spans multiple technical skills: coding, experimentation, and careful analysis. You will work in a tight feedback loop with the lab's researchers, turning the failure modes you find into stronger benchmark tasks.

This is a full-time W-2 employment position with Cincinnatus LLC, with the opportunity to be placed at a leading AI lab as part of their extended workforce. This role is fully remote within the United States, at approximately 35 hours per week.

2. Key Responsibilities

  • Probe models: Explore how frontier AI models behave on coding, ML, and analysis tasks, and find the spots where they quietly get things wrong.

  • Design challenges: Turn the weaknesses you find into well-crafted tasks that are hard for models but fair to grade.

  • Document findings: Write up what you discover clearly, with evidence and steps others can reproduce.

  • Strengthen tasks: Team up with task authors to close loopholes, shortcuts, and grading gaps.

  • Work as a team: Share insights with researchers and fellow experts so the benchmark keeps getting better.

3. Core Qualifications

  • MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy domain requiring data analysis and coding.

  • 1+ years of experience in a research, research-engineering, security, or AI-evaluation role.

  • Demonstrated ability to identify vulnerabilities, edge cases, or failure modes in LLMs or ML systems — through red teaming, adversarial testing, security research, or rigorous model evaluation.

  • Working proficiency in Python and Git, with the ability to script your own probes and analyses.

  • Strong familiarity with LLM capabilities, limitations, and evaluation techniques.

  • Past experience in AI training, model evaluation, or benchmark/task authoring is preferred.

  • A perfectionist mindset: high attention to detail, creativity in finding what others missed, strong written communication, and the ability to work independently through ambiguous, open-ended problems.

  • Ability to engage reliably for approximately 35 hours per week.

About Cincinnatus LLC

Cincinnatus LLC is an enterprise staffing company that partners with leading technology companies to source and employ highly skilled professionals for contingent and contract-based opportunities. Cincinnatus serves as the employer of record for these engagements, providing W-2 employment, payroll, benefits, and compliance, while placing employees directly within client teams to work on high-impact initiatives.

Roles hired through Cincinnatus are not project-based or freelance engagements. They are structured, role-based positions that typically involve part-time or full-time commitments, close collaboration with a client's internal teams, and integration into standard enterprise workflows.

Cincinnatus is a legal entity separate from Mercor. While opportunities may be discovered through Mercor's platform, employment, onboarding, payroll, and benefits for these roles are administered by Cincinnatus LLC.

Equal Employment Opportunity

Cincinnatus is proud to be an Equal Employment Opportunity employer. We do not discriminate based upon race, religion, color, national origin, sex (including pregnancy, childbirth, reproductive health decisions, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, genetic information, political views or activity, or any other legally protected characteristic.

Cincinnatus is committed to providing reasonable accommodations for qualified individuals with disabilities and disabled veterans throughout the job application process.

We consider all qualified applicants without regard to legally protected characteristics and provide reasonable accommodations upon request.

Earn up to $1,440 by referring

Share the referral link below, and earn up to $1,440 for each successful referral through this unique link. There's no limit on how many people you can refer. Restrictions may apply. Learn more
There's no limit on how many people you can refer. Restrictions may apply. Learn more
Don't know who to refer? Find relevant LinkedIn connections here.


One interview, real results
AI experts share how Mercor made hiring faster, fairer, and easier — with just one interview.

Posted 8 hours ago

$60-$90 / hr

Full-time · United States