NEWS / SEP.2026
Mercor removes OpenAI model evaluators for prohibited AI use, according to 404 Media
Mercor tells 404 Media that it removes contractors tasked with evaluating OpenAI’s models from its projects when it confirms prohibited AI use. The investigation published on September 22, 2026 describes removals and review guidelines stricter than Mercor’s general policy, which allows certain text edits.

According to 404 Media, Mercor removes OpenAI model evaluators for prohibited AI use
Mercor removes contractors tasked with evaluating OpenAI’s models for prohibited AI use, reports 404 Media in an investigation published on September 22, 2026. The company says it immediately removes experts from a project when it confirms such use. Mercor acts as an intermediary for these assignments; OpenAI is not their direct employer.
Joseph Cox interviewed three contractors working on projects linked to OpenAI, two of whom work for Mercor. Two anonymous sources describe removals. One contractor provided the journalist with what he presents as his termination letter, citing concerns about the authenticity of his work.
The dates of the removals and when the internal guidelines described took effect remain unknown. OpenAI declined to comment to 404 Media on the removals.
Edits are allowed, but evaluators must provide their own judgment
Humans evaluate the quality of models’ responses and write critiques intended to improve them. Project Lily, described by 404 Media on September 14, had already illustrated this work on ChatGPT, particularly to identify responses that are overly agreeable or portray the chatbot as a person.
In its public policy on language model use, Mercor allows corrections to grammar or word choice, as well as minor edits to phrasing or tone. The ideas and structure must come from the contractor.
The same policy prohibits using a model to evaluate the quality of other models’ responses or to produce the justifications for the work. It also prohibits asking a model to predict how code will behave or what its execution will produce. Mercor cites models’ lack of reliability in complex programming cases.
Correcting a sentence can help communicate human reasoning. Having a model produce the assessment of a response and its justification delegates the judgment for which the contractor was hired.
The internal documents described by 404 Media apply in particular to people who check other contractors’ work. For these reviews, they prohibit GPTZero, Grammarly and AI translation, including for writing comments or feedback. GPTZero is a tool for detecting AI-generated text. These guidelines impose additional restrictions beyond the minor corrections allowed by Mercor’s general policy.
Reviewers look in particular for repetition, certain uses of punctuation or very rapid completion, 404 Media reports.
“Model collapse” concerns recursive training
A study published in Nature on July 24, 2024 examines the recursive training of models on outputs from previous generations. When these generated data are used indiscriminately, the researchers observe degradation and a loss of diversity in the information learned. This phenomenon, called “model collapse,” concerns these training conditions, which are distinct from delegating an evaluation.
Mercor also provides for financial penalties
The public Project Offboarding page states that use that violates Mercor’s policies may result in removal from the project and withheld payment.
The Account Security and Conduct policy also lists possible penalties, including reductions or cancellations of payment for the work concerned, removal from projects and, in the most serious cases, suspension from the platform.