AI · Article 27 of 64

When a Tool Becomes an Oracle

When does rational reliance on a highly accurate AI become surrender of judgment?

The easy version of this question assumes a moment of surrender: someone hands over their judgment the way a driver hands over the wheel, and it happens all at once. What the research shows is stranger and more ordinary. Reliance usually grows without any decision that could be pointed to. Each individual act of trust is defensible on its own terms, and the accumulated result is a person who has quietly stopped checking.

The question this article works through is where that line sits: when does rational reliance on a highly accurate AI become surrender of judgment? The answer turns out not to be a threshold on accuracy. It depends on whether you are still capable of catching the error, whether you would find out if you could not, and whether the reliance is what you would choose if you could see the whole arrangement clearly. None of those conditions is about how good the model is.

The defect that accuracy does not fix

A tool that is right 90 percent of the time is worth using. That arithmetic is not in dispute. What the arithmetic hides is that a person’s job in an automated system is not simply to consume correct answers.

Consider what the operator is supposed to be doing. In monitoring tasks, they watch for the case the automation misses. In decision-support tasks, they bring context the system cannot see and apply judgment about cases that fall outside the training distribution. Both roles depend on the human remaining willing and able to disagree. That willingness is exactly what degrades under reliance, and it degrades in a way that ordinary accuracy figures do not capture.

This is the core of automation bias, and it was named in a 1999 experiment by Linda Skitka, Kathleen Mosier, and their colleagues. Participants performed a simulated flight task with a “very but not perfectly reliable” automated aid. The aided group did worse than an unaided control group at monitoring for events the automation did not flag. Two error types appeared. Errors of omission: the participant missed a problem because the aid had not prompted them to look. Errors of commission: the participant followed the aid’s recommendation even when it contradicted their own training and even when directly contradicting a fully valid indicator — the instrument said one thing, the aid said another, and some participants went with the aid.

An aid that is wrong ten percent of the time should improve performance by being right ninety percent of the time and cost something by being wrong the rest. It did something else: it made people stop looking. The cost was not localized to the errors. It was spread across the whole task.

Two mechanisms, one bottleneck

Automation bias is often discussed alongside automation complacency, and the two are worth separating because they respond to different things.

Complacency is a monitoring failure under load. When several tasks compete for attention, and one of them is being handled by a reliable machine, attention flows elsewhere. The operator is not fooled about the automation’s reliability; they are simply not watching. Complacency appears in both novices and experts and is not fixed by practice — multi-task load keeps pushing attention away from the boring, usually-correct display.

Bias is a judgment failure. The recommendation carries weight beyond what its reliability justifies, and the operator’s own assessment gets anchored to it or abandoned entirely.

In a widely cited 2010 review in Human Factors, Raja Parasuraman and Dietrich Manzey argued that the two share an underlying constraint: attention (Parasuraman & Manzey, 2010). Their synthesis is uncomfortable reading for anyone hoping to solve this with better instructions. Automation bias “occurs in both naive and expert participants, cannot be prevented by training or instructions, and can affect decision making in individuals as well as in teams.”

That claim deserves to be read precisely. It does not say training is useless. It says that telling people to be vigilant is not sufficient, because vigilance is the resource the setup is spending. Awareness of automation bias is necessary and unhelpfully cheap.

What a decade of review evidence added

The empirical literature is now large enough to be summarized rather than sampled. A systematic review by Kate Goddard, Abdul Roudsari, and Jeremy Wyatt screened more than 13,000 papers and included 74 in a structured analysis of automation bias across domains (Goddard et al., 2012). The pattern that emerged was consistent: overreliance shows up wherever a system offers advice on a decision the human remains nominally responsible for.

Their analysis is useful for what it says about mediators — the conditions that make bias worse or better. Cognitive style, domain experience, trust in the system, workload, task complexity, and time pressure all moved the effect. Notice the shape of that list. Most of it is not about the AI.

The medical findings are the ones with the sharpest stakes, because there the errors land on patients. A follow-up study by the same authors put 26 UK general practitioners through twenty prescribing scenarios, with the simulated decision support giving wrong advice in six of them and operating at a stated 70 percent reliability (Goddard et al., 2014). The participants changed their prescription in 22.5 percent of cases. Pre-advice accuracy was 50.38 percent and rose to 58.27 percent after the advice — a net gain of about 8 percentage points. The gain, though, was assembled out of two opposing movements. In 5.2 percent of all cases, a clinician who had been right switched to a wrong answer. The aid was worth using, and it manufactured new errors at a rate low enough to be invisible in the aggregate.

The study’s other finding is the one that ties back to the attention argument. Lower clinical experience was associated with more switching. The clinicians most willing to be corrected by the system were the ones with the least independent basis for knowing when the correction was wrong.

Physicians, advice, and the label on it

The cleanest test of the attribution question came from a 2021 study by Susanne Gaube and colleagues in npj Digital Medicine (Gaube et al., 2021). Two hundred sixty-five physicians — 138 radiologists and 127 internists and emergency physicians — diagnosed cases with decision support. Here is the design detail that matters: all the advice was human-generated. The researchers simply relabeled some of it as coming from an AI system and some as coming from a human expert.

The results complicate both the optimistic and the alarmist stories.

When the advice was wrong, diagnostic accuracy fell — regardless of whether the label said AI or human. Trust was not misallocated to machines specifically; it was misallocated to advice. Then the groups diverged. Among internists and emergency physicians, 41.73 percent always followed the incorrect advice. Among radiologists, 27.54 percent did. Radiologists also rated the AI-labeled advice lower than the human-labeled advice, which the authors describe as algorithmic aversion — the same recommendation discounted when it arrived from a machine.

Two lessons sit in that split. First, the susceptibility to advice is not really about AI; it is about what happens to a professional’s independent assessment when a confident recommendation arrives first. Second, the AI-specific component can run either direction. Some users over-trust the machine; some discount it. Which way a particular person leans is a property of the person and the context, and a system designed on the assumption of one bias will misfit half its users.

Explanations do not fix trust, and sometimes worsen it

The intuitive fix for overreliance is to show your work. Explain the recommendation, surface a confidence score, teach the user to calibrate.

A 2021 study by Zana Buçinca, Maja Barbara Malaya, and Krzysztof Gajos tested this directly, with 199 participants making decisions with AI assistance (Buçinca et al., 2021). They compared plain AI assistance against assistance with explanations, and against cognitive forcing functions — designs that make it harder to accept a suggestion without thinking. The forcing conditions included committing to your own answer before seeing the AI’s, deliberately delaying the decision, and making the AI’s suggestion available only on request.

The findings are worth stating baldly. Explanations alone did not reduce overreliance; in some conditions they increased it, apparently by making the recommendation feel more credible. The cognitive forcing functions did reduce overreliance. And they were the conditions participants liked least.

That last clause is the one that should govern design. Measures that protect judgment are experienced as friction. A product team optimizing for satisfaction and speed will, without any ill intent, select against them. The study also found a distributional effect worth naming: the interventions helped participants with a higher measured need for cognition more than others, which the authors frame as the risk of intervention-generated inequality. A safeguard that only works for people already inclined to deliberate is not much of a safeguard.

When the combination is worse than either part

It would be reasonable to hope that these effects are offset by the gains from combining human and machine. A 2024 systematic review and meta-analysis tested that hope at scale. Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone pooled 106 experiments reporting 370 effect sizes, each comparing human alone, AI alone, and human plus AI (Vaccaro et al., 2024).

Three results matter here.

On average, the human-AI combinations performed significantly worse than whichever of the two did better alone. The effect was small (Hedges’ g = −0.23) but consistent, and it means the collaboration was often not worth its overhead. The combinations did beat the human alone — so AI helps people — but they did not beat the AI alone, which raises an obvious question about whether the human added anything.

Task type split the results sharply. On decision tasks, the combination did worse. On creation tasks, it did better. The authors’ explanation is that creation has a natural division of labor — human judgment about what is good, machine throughput on the mechanical parts — while a decision is a single indivisible act that two parties cannot each half-perform.

The most consequential moderator was relative skill. Where humans outperformed the AI alone, the combination gained. Where the AI outperformed humans alone, the combination lost.

Read that alongside the automation-bias literature and the pieces fit. The case where the human most needs to defer is the case where their deference is most costly, because there is no residual signal left for them to contribute and no mechanism keeping their judgment calibrated. Conversely, the combination works when the human is genuinely the better performer and the AI is genuinely subordinate — a division of labor rather than a handoff.

The meta-analysis also found that neither explanations nor confidence levels significantly moderated the results. That is the third independent line of evidence against the intuitive fix.

What the survey evidence says about everyday work

The studies above are laboratory tasks with a right answer. A 2025 CHI paper by Hao-Ping Lee and colleagues looked at what is happening in ordinary knowledge work, surveying 319 knowledge workers about 936 specific instances of using generative AI on the job (Lee et al., 2025).

Two correlations stand out. Workers who reported higher confidence in GenAI reported engaging in less critical thinking. Workers with higher confidence in themselves reported more. And the character of the thinking shifted rather than disappearing: less generation from scratch, more verification, integration, and stewardship of AI output — checking sources, fitting output into existing context, deciding whether to use it at all.

That is a more nuanced finding than either the alarm or the reassurance. The work changes shape. But the survey also identifies something the automation literature predicts: when the human’s role becomes verification, the quality of the verification depends on whether they have the knowledge to verify. A worker who never developed the underlying skill is being asked to audit output in a domain where they have no independent basis for judgment.

The regulatory answer, and its limits

This problem is no longer purely academic. The EU AI Act’s Article 14 makes human oversight a design requirement for high-risk systems, and its text names automation bias explicitly (EU AI Act, Article 14). Overseers must be able to understand the system’s limits and monitor for anomalies; to remain aware of the tendency to over-rely on output; to interpret output correctly; to decide not to use it, disregard it, override it, or reverse it; and to intervene or stop the system through a stop procedure.

The framing is notable. Oversight is defined as a set of capabilities the operator must actually have, not a checkbox that a human signed off. The override must be real, the stop must exist and work, and the operator must know the failure modes well enough to notice one. For remote biometric identification, the Act goes further and requires separate confirmation by two competent people before any action is taken on an identification.

Two limitations deserve stating plainly. First, the requirement binds high-risk systems and leaves most everyday AI use — the assistant, the writing tool, the coding helper — outside its reach, which is where a great deal of the overreliance now happens. Second, a legally mandated override capability does not create the judgment to use it. A stop button that no one presses because pressure, workload, and trust all point the other way satisfies the letter of the requirement and none of its purpose. Goddard and colleagues found that training and instructions are not sufficient; a regulation that supplies instructions and capabilities has supplied the ingredients, not the outcome.

Where the line actually falls

The temptation is to look for an accuracy threshold: above some level of reliability, deference is rational and the problem dissolves. The evidence does not support that shape.

Take a system that is right far more often than the human. Deferring every time maximizes the expected number of correct answers on today’s decisions, and at the same time destroys the human’s ability to detect the one case per thousand where the system is confidently wrong and the stakes are high. Expected value on the task and capability to supervise the system point in opposite directions for the same person in the same moment. That is why “rational reliance” and “surrender of judgment” are not separated by a line in accuracy space. They are separated by whether the reliance is recoverable.

Four questions distinguish them.

Can you detect the failure? If you have no independent basis for judging output — no domain knowledge, no ability to check a source, no sense of what the answer should look like — then your oversight is nominal. The human in the loop is a signature, not a judgment.

Would you find out? A system can be accurate, and wrong in ways that never surface. If errors are logged, attributed, and fed back, trust can recalibrate. If wrong answers simply disappear into the world, trust can only drift upward.

Is the reliance reversible? A person who can work without the system, who retains the skill and the habit, is using a tool. A person who has never done the task unaided, or who would be unable to produce a defensible answer without it, occupies a different position. The distinction is not about nostalgia for doing things the hard way. It is about whether removing the system removes the capability.

Is this what you would choose for yourself? The question is whether you were shown the arrangement — the accuracy, the failure modes, the override, the fact that your skill will be shaped by how you use it — and chose it anyway. The automation-bias literature describes people who would not choose to stop checking but who stop checking because attention is finite and the machine is usually right. Consent given in ignorance of how the arrangement will change attention is thin consent.

Designing for the thing you actually want

If the goal is to keep judgment alive while capturing the accuracy benefit, the design implications follow from the evidence rather than from intuition.

Make disagreement possible and low-cost. The override has to exist, be discoverable, and carry no penalty — social or economic. Buçinca and colleagues showed that mechanisms which force a moment of independent judgment work, and they work even though people dislike them. The dislike is the price, and it argues for making the friction internal to the workflow rather than optional.

Surface uncertainty where it can be acted on. Not a decorative confidence bar. A system that can say “this is a case I am poorly calibrated on, and here is what you should check” is doing something functional. Note that the meta-analysis found generic confidence displays had no measurable effect; vagueness in the display is part of why.

Give people the basis to verify. Verification skills are not free and do not survive disuse. A domain where everyone has delegated the underlying task has no reserve of judgment to audit with, which is the deepest form of the problem and the slowest to appear.

Watch the distributions. Interventions help the deliberative more than the impulsive, the expert more than the novice. A safeguard that only works for people already inclined to use it accrues its benefits to those who need them least, and that pattern should be measured rather than assumed away.

Instrument the overrides. Override rate, time to intervention, and rate of disagreement with flagged cases are measurable. A service that does not know how often its humans disagree with it cannot tell calibrated deference from drift.

The part that is still open

The literature is clear about the mechanism and less clear about the magnitude in the wild. Laboratory tasks have right answers and short horizons. A knowledge worker using a writing assistant for a year is in a situation no experiment has fully modeled, because the outcome accrues slowly, is difficult to score, and interacts with a career’s worth of accumulated skill.

The strongest defensible summary is this. Highly accurate AI genuinely improves average performance, and the meta-analysis shows combinations beating the human alone. It also shows combinations failing to beat the better of the two, and losing specifically on decision tasks and specifically when the AI alone is better. Automation bias is real, it appears in experts, and it is not removed by training or by explanations. The legal system has begun to require oversight without being able to require the judgment that makes oversight work.

That combination points to an asymmetry worth carrying out of this one. Reliance on a tool you could out-argue is still reliance. Reliance on a tool you could no longer check is something else, and the transition between them is not marked by any event — it is marked by a skill you can no longer exercise and an override you no longer reach for.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home