AI · Article 61 of 64

From Humans Using Machines to Human-Machine Systems

At what point does it stop making sense to measure the intelligence of the human and machine separately?

In 2024, a team led by Michelle Vaccaro, Abdullah Almaatouq, and Thomas Malone at MIT published a systematic review and meta-analysis of 106 experiments. Each experiment compared three arrangements on the same task: people working alone, AI systems working alone, and people and AI working together. The expectation most of us carry — that the combination beats both — did not survive contact with the data. Across 370 effect sizes, human-AI combinations performed slightly worse than the better of the two components on its own. The combinations did, on average, beat people working unaided. But the pattern that matters is narrower than either headline: teams gained on content-creation tasks and lost on decision tasks, and the direction of the loss tracked a simple variable. When the human alone outperformed the AI alone, collaboration tended to help. When the AI alone outperformed the human, collaboration tended to hurt (Vaccaro, Almaatouq, and Malone, 2024).

That finding is a useful place to start because it says something structural rather than anecdotal. Whether a human-machine arrangement behaves like a system — a thing whose joint capability exceeds that of its parts — depends on the interface between the two, not on either side’s score. A better model can make the pair worse. A worse model can sometimes make it better. The relationship between component quality and combined quality is not monotonic, and that fact is the reason this article’s question is harder than it first appears.

“At what point does it stop making sense to measure the intelligence of the human and machine separately?” sounds like a question about how smart machines eventually become. It is really two questions wearing one coat. The first is empirical: under what conditions do coupled arrangements produce something neither part produces alone, and can that be predicted from measurements of the parts? The second is conceptual: even when the parts can be measured, is “the intelligence of the human” or “the intelligence of the machine” the correct unit of analysis in the first place?

Complementarity is a property of the pair, not the parts

The most careful formal treatment of the first question comes from Mark Steyvers and colleagues, who built a Bayesian model of what they call human-AI complementarity (Steyvers et al., 2022). Their central result is easy to state and easy to underestimate: complementarity is achievable even when the human and the AI differ substantially in accuracy, provided their errors are decorrelated in the right way. The model yields a bound on how different two decision-makers’ accuracy can be before combining them stops helping. Critically, that bound is set by the correlation between their confidence signals, not by their skill.

Two consequences follow. First, you cannot read off team performance from individual accuracies. Two components at the same accuracy can form a strong pair or a useless one depending on whether they are wrong at the same times, and the wrong-at-the-same-times structure is invisible to any benchmark that scores each side in isolation. Second, the human’s contribution is often information about the human’s own uncertainty. Steyvers’s group found that eliciting a person’s confidence in their judgment — asking them, in effect, to report how sure they are — measurably improved the combined system’s performance. The human was adding value not by being more accurate, but by supplying a signal the machine did not have. A system that treats the human as merely another classifier misses this.

The design literature reaches the same conclusion from the opposite direction. In a study of prediction tasks with varying levels of AI accuracy, researchers found that a less accurate AI could function as a better collaborator than a more accurate one, because the less accurate system’s errors were positioned where the human could catch them, while the more accurate system’s errors landed in exactly the region where the human had already stopped paying attention (Inkpen et al., 2023). They also found that the outcome depended heavily on whether people had domain expertise and on how the automation was tuned for reliance. This is a genuinely awkward result for anyone who assumes that better components make better systems. It suggests that the properties worth measuring live at the boundary: where errors land, whether the human’s attention is directed at the right places, and whether the interface preserves the conditions under which the human’s judgment still applies.

The offloading is not automatically the problem

Before treating dependence as a pathology, it is worth remembering that offloading cognition is how human beings got good at anything. Writing is offloaded memory. Arithmetic notation is offloaded calculation. Maps, calendars, checklists, and consulting a colleague are all ways of pushing work out of a single skull and onto a substrate that does it more reliably.

Betsy Sparrow, Jenny Liu, and Daniel Wegner documented a clean version of this in 2011. When people expected to have future access to a piece of information, they remembered the information itself less well — but they remembered where to find it better (Sparrow, Liu, and Wegner, 2011). The paper has been widely misread as evidence that search engines are damaging memory. It is more accurately a demonstration that memory allocates itself according to what is likely to be available, which is an adaptive strategy and roughly what a well-functioning person should do in a world full of accessible information.

Samuel Risko and Sam Gilbert formalized this with a metacognitive framework for cognitive offloading across domains (Risko and Gilbert, 2016). Their review finds that offloading generally improves performance and that people use it sensibly a good deal of the time. But it also identifies the specific failure mode that matters here: offloading decisions depend on metacognitive judgments about one’s own abilities and about the reliability of the tool, and those judgments can be systematically wrong. A person can persist in offloading after the tool has started failing, or refuse to offload when it would obviously help. The risk is not delegation as such. It is delegation that has lost its check.

That reframing makes the question sharper. The interesting question is not whether the human still remembers things without AI. It is whether the human’s metacognitive layer — the part that decides when to trust, when to verify, and when to override — is still calibrated to reality. Everything about human-machine systems reduces to whether that layer survives contact with a tool designed to be fluent, confident, and usually right.

What the working evidence says about that layer

A 2025 study by Hao-Ping (Hank) Lee and colleagues gives the most concrete field evidence to date. They surveyed 319 knowledge workers and analyzed 936 real examples of generative-AI use at work, then coded what the workers actually did with the output (Lee et al., CHI 2025). Two findings stand out.

The first is a negative correlation between confidence in the AI and the effort spent thinking critically about its output. Workers who trusted the tool most engaged in the least verification. The second is the shape of the work that remained: as the AI absorbed generation, the human’s role moved toward verifying claims, integrating responses into real workflows, and taking responsibility for the result — what the authors call a shift toward stewardship. They also connect the pattern to the classic automation literature, where the human in a highly automated loop is expected to monitor a system that almost never fails and to intervene correctly on the rare occasion that it does, which is a notoriously difficult task for human attention. The failure mode of a very good tool is not that it errs often. It is that it trains its user to stop looking.

This is where the two halves of the question start to merge. If the human’s contribution to the pair is increasingly verification and stewardship, then the human’s measurable skill on the original task may decline while the pair’s performance stays high — and possibly keeps improving. Measuring the human alone would then understate what the pair can do. Measuring the pair alone would overstate what the human can do. Both numbers would be true and neither would be the right number, because the operative causal unit is the loop, and the loop is not the same thing as either endpoint.

Where separate measurement comes apart

Consider a concrete case. An analyst uses a model to draft a regulatory risk assessment, checks the citations, rewrites two sections, and signs it. Where did the judgment live? The model supplied recall, structure, and alternatives. The analyst supplied the choice of what mattered, the recognition that one citation was being over-read, and the willingness to attach their name to the result. If you administered the analyst a closed-book test of the same facts, they might score worse than the model. If you gave the model the task alone, it might produce a fluent document with a plausible citation error that no one caught. Neither isolated measurement reflects the process that actually occurred, and the process is the thing that produced the outcome.

This is not a mystical claim about merged minds. It is a mundane claim about which variables predict outcomes. Practically every consequential human-machine arrangement already has this structure — a surgeon with imaging, a pilot with autopilot, a radiologist with a detector — and in all of them the interesting quantities are properties of the coupling: who has the information, who can override, how quickly the human notices a divergence, whether the human’s attention is trained to the right places. When those properties are good, the pair outperforms both. When they are bad, the pair underperforms the better component, which is exactly what the 106 studies found on average.

There is a temptation to conclude that as models improve, the human role shrinks toward nothing and the boundary eventually dissolves. That conclusion does not follow. The meta-analysis found that collaboration helped when humans were better than the AI on the task and hurt when they were worse, which means the binding variable is comparative competence and error structure, not absolute machine capability. A more capable model can widen or narrow the human’s useful contribution depending on how the interface routes the human’s attention. The direction is a design choice. That is the part usually left out of predictions about machines eclipsing people: capability improves on a curve, but where the human’s comparative advantage lands is determined by whether anyone bothered to build the coupling so the advantage still has a place to sit.

Reversible coupling as a design principle

If the unit of analysis is the pair, then the properties worth engineering are properties of the coupling, and the most important of them is reversibility. A well-built human-machine system should let the human leave. That means the memory the system accumulates must be exportable; the competence the human builds must be theirs and not rented session by session; and the failure of the service must degrade the pair rather than disable the person. These are engineering decisions, and the same underlying capability can be coupled reverently or extractively depending on how they are made.

The asymmetry of exit is where the governance stakes are sharpest. A person can stop using a tool; a provider can revoke access to the tool, and with it the memory, workflow, and accumulated personalization built inside it. When one vendor holds a person’s notes, calendar, correspondence, and now their reasoning traces, the practical difference between dependence and ownership is the difference between a product they can leave and an arrangement they cannot. This is not an argument against ambitious tooling. It is an argument that portability and local fallback are load-bearing features rather than compliance pleasantries, because they determine whether the human in the pair is a participant with standing or a component that can be deactivated.

There is a corresponding obligation on the human side, and it is the one that gets skipped. Reversible coupling is worthless if the human no longer has the metacognitive habits to exercise the option. Sparrow’s result implies that people will remember less of what a system will remember for them, which is fine, and Risko and Gilbert’s implies that people will misjudge when to offload, which is not. The only durable safeguard against the second is the one that looks least like a technology policy: people who can still do the underlying work well enough to tell when the machine is wrong, and systems whose design keeps them exercised rather than rusted.

What would actually tell us the boundary has dissolved

Claims that human and machine have become one system are cheap, so it is worth saying what would make the claim true rather than evocative. Four observations would count, and they are all measurable.

First, the pair would need to beat the best component not in one lab task but across many domains with different error structures, since the current evidence is that gains concentrate in creation tasks and losses concentrate in decision tasks. Second, the human’s contribution would need to stop being characterizable as supplying a value or a preference — as it is in the Steyvers result, where the human’s main addition is confidence information — and start being genuinely generative in a way no benchmark of the machine can capture. Third, removing the machine would need to reveal not a hollowed-out person but a still-competent one who is now measurably slower, which would distinguish healthy division of labor from atrophy. Fourth, the human would need to remain able to inspect and override the output at a level of detail that makes the override real rather than theatrical.

Today’s evidence supports a partial version of the first and third, is genuinely unproven on the second, and is worse than people assume on the fourth. And there is a fifth requirement that runs the other way, which no amount of engineering solves: whatever values the pair is optimizing still have to be chosen by someone with the standing to choose them. An intelligence can compute the consequences of a policy and summarize the traditions that bear on it. It cannot, by being more capable, make a genuinely contested moral question settled — the authority to decide is not a form of capability, and it does not transfer along with the computation. That is the sense in which the boundary between human and machine should hold even where measurement stops distinguishing them: the pair can be a single performance unit while remaining two distinct sources of authority.

The question that replaces the old one

Start with the finding that made the separate-measurement frame wobble: in most of the experiments anyone has run, the combination is not automatically better than its best part. The conditions for a real system are specific, they involve error structure and attention rather than raw scores, and they can be engineered either way. That means the useful question is not “is the machine as smart as the human yet” or “when do the two merge.” It is “what is this particular coupling for, does it beat its parts, and can the human leave.”

Measured that way, the boundary stops being a line that dissolves at some threshold of capability and becomes a set of design properties that can be maintained or allowed to rot. The intelligence worth tracking is not the intelligence of the human or of the machine. It is the intelligence of the arrangement, plus the separate, indispensable question of whether the person inside it is still capable of walking away and still competent enough that walking away would mean something. That pair of measurements is more awkward than a single number and considerably more informative, which is usually the sign that you have found the right unit of analysis.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home