Humans Monitoring AI: New Study Reveals Fatal Flaw in Our Strategy

14 September 2026

Imagine a physician reviewing a colleague’s diagnosis and deeming it deficient. Now imagine that same diagnosis labeled “generated by artificial intelligence.” The question pursued by Sofoklis Goulas, Rigissa Megalokonomou, and Panagiotis Sotirakopoulos was uncomfortable: does that detail change anything? Is the expert’s judgment truly impartial in the face of the machine?

The answer, published recently in the scientific journal PNAS Nexus, is decisive: no. And the margin of difference is by no means negligible. Far from accepting the algorithmic verdict with the same deference as that of a colleague, the fact that they are auditing the work of a machine seems to undermine the authority of the evaluation in the eyes of the expert, rendering the expert quiet and unconsciously more interventionist when faced with a criterion that is too harsh to begin with.

The experiment that measured the gap

The team designed a preregistered randomized experiment in which education-education experts graded the same student work, with one key variation: one group was told that the grading recommendations came from an AI algorithm; the other, that they came from a human grader. The content was identical. Only the label changed.

The experts who believed they were supervising an AI showed an equity gap 22% larger than those who believed they were supervising a human. In practical terms, this means that when the AI was perceived as harsher than the expert would have been, the expert tended to compensate for the student by adjusting upward. But when the label was “human,” the same level of severity produced a smaller correction.

“The bias does not go in the direction most people imagine. It’s not that we blindly trust AI: it’s that we apply a different yardstick of authority. When a human is harsh in evaluation, the expert tends to respect that severity; but when the machine shows the same hardness, the expert deauthorizes it and intervenes to compensate the student.”

The finding inverts the conventional narrative. It is not that experts are hostile to algorithmic recommendations. The problem is more subtle: the label “AI” unconsciously triggers a different corrective frame.

Why the “the human in the loop” isn’t enough

The concept of human-in-the-loop, the idea that placing a person between the AI’s decision and its consequence is all that’s needed to stop automation’s mistakes, has long been pitched as the great firewall. If the AI errs, the human detects and corrects it. The Goulas and colleagues’ experiment shows that this firewall can have systematic leaks.

The mechanism the researchers propose relates to how we allocate authority and agency. When the decision comes from a person, there is implicit respect for their professional judgment. When it comes from a system, that presumption of authority dissipates. The result is that the supervising expert corrects the AI more actively, altering the evaluation standard simply because they feel more legitimate to retract a machine’s punishment than a human peer’s.

This has implications that extend far beyond exam scoring. In any setting where a professional reviews an AI system’s output—whether a radiologist verifying a diagnosis assisted by AI, or a judge evaluating a sentencing recommendation—the bias could be operating in a quiet way, distorting precisely the step that was designed to neutralize the algorithm’s errors.

Researchers label this phenomenon the grading fairness gap: the gap between the standard of correction we apply to the machine and the one we apply to a peer.

The paradox of the vigilant expert

There is something deeply paradoxical in this result. Those who supervise AI are often the most highly qualified in their field. They are theoretically best equipped to detect an error. And yet, the same label that should trigger their skepticism appears, according to this experiment, to relax their judgment in certain contexts.

The researchers do not attribute the effect to incompetence or cognitive laziness, but to an implicit adjustment mechanism. The expert doesn’t “let their guard down” consciously. What happens is that the frame of reference shifts: the AI is perceived as a tool that can operate in ranges different from those of a human reviewer, and that perception alters the threshold at which the supervisor considers a correction necessary.

“The problem isn’t that humans fail to monitor AI. The problem is that they don’t know when they’re failing.”

This point is the study’s most unsettling aspect. An error that the expert does not realize they are making is an error that they cannot correct. The solution, then, isn’t simply to add more layers of human oversight without any other changes: it requires understanding what kinds of biases the label itself triggers in the supervisor.

Not all situations are equal

The experiment occurred in an educational assessment context, with experts whose task was to review students’ grades. It is a controlled, well-defined environment. The authors do not claim that this bias operates with the same strength in life-or-death contexts, such as emergency medicine, autonomous driving, or criminal justice, where the consequences could trigger very different attention dynamics.

What they do offer, however, is robust empirical evidence, with randomized, preregistered design, that bias exists and is measurable in real-world expert workloads. That is enough to challenge the assumption that “putting a human” in the loop resolves the problem of algorithmic errors.

The open question, and what the researchers themselves point to as the next step, is whether the effect varies by type of decision, by the supervisor’s level of specific AI-system training, or by the visibility of the error in question. Not all AI errors are equally detectable, and not all experts have the same degree of algorithmic literacy.

The next problem

If the bias exists and is unconscious, how can we fix it? The most obvious—and uncomfortable—answer is that the solution cannot depend solely on the expert themselves. We need auditing protocols that do not reveal the source of the recommendation until the supervisor has issued their own judgment, or double-blind systems that remove the label as a contaminating variable.

Goulas, Megalokonomou, and Sotirakopoulos do not propose a definitive cure. They argue, instead, that the problem should be taken seriously before deploying human-supervision systems as a corrective guarantee in high-risk environments. Because if the human in the loop does not see what they should see, the loop does not close.

What remains to be known is whether this bias is trainable, whether it can be attenuated through targeted training, or whether it is so deeply embedded in how we assign agency to machines that it will resist any deliberate corrective protocol. That is the question the study has just placed on the table.

Olivia Parker

I write about the trends, stories and cultural shifts that catch my attention, from everyday discoveries to unexpected ideas from around the world. Based in Flin Flon, I’m always looking for the next story worth remembering.