AI · Article 52 of 64

AI as a Cognitive Exoskeleton

What is the healthiest metaphor for a personal AI: replacement brain or cognitive exoskeleton?

J. C. R. Licklider’s 1960 essay “Man-Computer Symbiosis” sketched a division of labor that has aged better than the hardware it described. People would set the goals, frame the hypotheses, and define the criteria for success; computers would absorb the routinizable work that prepares the way for insight (Licklider, 1960). The machine would carry load. The person would choose the destination.

That is an exoskeleton, described decades before anyone applied the word to thinking. Licklider was writing about interactive computing and time-sharing, and the shape of his proposal is the shape of a wearable frame: an apparatus that removes weight from the person inside it while leaving that person responsible for where the weight is going.

A competing metaphor has since crowded in. In the replacement-brain picture, a personal AI is a superior mind that a person consults and then defers to. Cognition gets outsourced the way a firm outsources payroll. In the exoskeleton picture, the same underlying model is bound to a person’s own intentions and extends their reach. Both pictures can be assembled from identical model weights, and the difference between them has little to do with how clever the software is. It has to do with where authority sits when the person and the system disagree.

Whether the exoskeleton is the healthier metaphor is partly an empirical question and partly a design commitment. The empirical half asks what actually happens to human competence when people lean on these tools. The design half asks what the tools are permitted to do on their own authority. This essay treats the first as evidence, the second as engineering, and keeps a third category, speculation, clearly marked.

Two metaphors pull in different directions

An exoskeleton has a defining property: it is worn. The wearer can remove it. The frame senses force and supplies assistance, and the human gait stays the reference signal. Gait-rehabilitation exoskeletons are generally built around this idea, supplying help that scales with the movement the user is attempting rather than following a fixed schedule. The assistance is large when the person is struggling and small when the person is not, and the person stays responsible for balance and direction throughout.

A replacement brain has a different defining property: it is consulted. A question goes in and an answer comes out, and the internal work that would have produced the answer never happens.

The two arrangements differ along three axes. On memory, an exoskeleton augments retrieval while a replacement brain bypasses it. On initiative, the exoskeleton leaves the generation of goals with the person, whereas a replacement brain increasingly supplies the goals along with the answers. On failure, removing the exoskeleton leaves a slower competence intact, while removing the replacement brain reveals that the task never belonged to the person in the first place. One arrangement produces someone who is out of practice. The other produces someone who was never in practice.

Because so much writing in this area slides between evidentiary categories, the categories are worth naming once. A demonstrated claim comes from a peer-reviewed study with identified participants and stated conditions. A plausible-engineering claim assembles demonstrated parts into something that has not been built or tested at the claimed scale. A speculative claim reaches past every experiment currently on record, and it tends to be marked by confident language about minds rather than modest language about benchmarks.

What the evidence actually shows about offloading

The oldest and most replicated finding here is that people treat networked information as transactive memory. In four studies published in Science in 2011, Betsy Sparrow, Jenny Liu, and Daniel Wegner found that when participants expected to be able to look something up later, they remembered the fact itself less well and remembered where to find it better (Sparrow, Liu, & Wegner, 2011). The result describes a reallocation rather than an injury: the mind stored an index instead of the contents. People have always distributed memory across spouses, colleagues, and libraries, and the 2011 finding showed that the distribution extends to a machine, which becomes part of the memory system rather than staying beside it.

Navigation offers a sharper case, because a spatial skill can be tested directly. Véronique Bohbot’s group at McGill measured lifetime GPS use in fifty regular drivers and then tested spatial memory in virtual mazes. Participants who relied on GPS more performed worse on hippocampus-dependent spatial memory when navigating without it, and in a three-year follow-up with thirteen of the original participants, greater GPS use was associated with a steeper decline (Dahmani & Bohbot, 2020). The researchers checked the obvious alternative account and found that heavy GPS users had not reported a poorer sense of direction at the outset, which weakens the idea that people with weak spatial memory simply select themselves into GPS. The main result is cross-sectional and the longitudinal arm is small, so the causal direction stays open to argument. Read conservatively, it remains the cleanest published example of a cognitive skill that appears to weaken as a tool absorbs the underlying work.

The generative-AI wave has produced a cluster of studies with a common shape. Hao-Ping Lee and colleagues at Microsoft Research and Carnegie Mellon surveyed 319 knowledge workers about 936 real work tasks and found that greater confidence in generative AI predicted less critical thinking, while greater confidence in oneself predicted more (Lee et al., 2025). The mechanism they describe is a migration from producing work to supervising it: verification, integration, and stewardship replace generation. That migration is the exoskeleton effect in miniature, though the survey measures self-reported behavior and cannot show what skill remains.

A 2025 survey of 666 participants in Societies found a negative correlation between frequent AI tool use and a measure of critical thinking, mediated by self-reported cognitive offloading (Gerlich, 2025). The paper drew wide attention and also a published correction. Its design is correlational and its outcome measure is self-report, so it cannot establish that AI use caused a decline.

The most-cited study in this cluster is a preprint from the MIT Media Lab. Nataliya Kosmyna and colleagues recorded EEG while participants wrote essays with a large language model, with a search engine, or with no tools at all (Kosmyna et al., 2025). Across three sessions, the language-model group showed the weakest brain connectivity, reported the lowest sense of ownership over their writing, and had more difficulty quoting their own work. It is a preprint that has not completed peer review; fifty-four participants finished the first three sessions and eighteen finished a fourth; and the authors themselves cautioned against reading the results as evidence that AI damages the brain. The fair summary is that a small, unreviewed study found a pattern worth following, and that the pattern is consistent with the older findings on transactive memory and GPS.

Pulled together, the evidence supports a narrow claim. When a tool absorbs routine cognitive work, measurable performance on that specific work can decline, and a person’s sense of authorship can thin along with it. The evidence does not support the claim that using an AI makes a person broadly less intelligent, and it does not support the opposite claim that heavy use carries no cognitive cost.

Where genuine amplification has been demonstrated

The strongest clinical demonstration of cognitive amplification is not a chatbot. It is a neuroprosthesis. In 2023, a team at the University of California, San Francisco reported that a participant with severe paralysis from a brainstem stroke, unable to speak or type, could produce text from attempted silent speech at a median of 78 words per minute with a median word error rate of 25 percent, decoded from a 253-channel electrode array implanted over her speech cortex (Metzger et al., 2023). The system also synthesized audible speech in a voice personalized to her pre-injury voice and drove a facial avatar. The result is peer-reviewed and remarkable: roughly five times the rate she reached with the head-tracking assistive device she had been using.

The conditions matter as much as the result. One participant, in an invasive clinical trial, with decoding calibrated to her own cortex. The authors presented the work as a proof of concept for a medical device. The distance between it and a consumer product is measured in years of safety testing, surgical risk, and cost.

Most of what a reader will meet under the heading of a cognitive exoskeleton sits well below that peak. Retrieval, summarization, translation, and drafting are demonstrated and widely deployed. Persistent memory across sessions is plausible engineering with uneven reliability, and the honest test is what happens when a stored fact is wrong and correcting it matters. Multi-step agents that plan and execute open-ended research are early engineering whose results depend on how tightly the task is scoped and how errors are caught. Claims about directly coupled cognition are speculation.

The benefits at the lower rungs are not small. Lowering the cost of retrieval and synthesis changes who can attempt a difficult problem. A clinician can survey a literature in an afternoon. An independent inventor can run an analysis that once required a laboratory. A student who struggles with print can hear a text in plain language. None of these requires the system to be autonomous, and none requires it to be correct about everything. They require speed, tolerable reliability, and inspectability.

A complete accounting has to price dependence that is purely commercial. If a person’s drafts, notes, and working context live inside one vendor’s product, then the exoskeleton has an owner. Switching costs become a form of lock-in, and the terms on which the apparatus can be withdrawn or repriced belong to someone else. Portability is the unglamorous design requirement that decides whether the metaphor holds. An apparatus you can remove is an apparatus. One you cannot is a limb.

Substitution has a signature

Substitution rarely arrives announced. It shows up as convenience, and it has a recognizable signature in five parts.

The first is skill atrophy through disuse. The GPS and transactive-memory results point to an ordinary mechanism: a capacity that goes unexercised does not hold its level. This is demonstrated in the narrow cases that have been studied and plausible as a broader pattern.

The second is the movement from doing to supervising without a matching gain in the ability to supervise. When the production is delegated and the oversight is also assisted, the person in the loop may be approving output they cannot independently evaluate. The survey evidence describes the shift toward verification and stewardship; whether verification actually sharpens with practice is an open question rather than a settled result.

The third is capture of the objective. An exoskeleton is only as good as the direction it amplifies. A system tuned for engagement, throughput, or revenue will find the shortest route to its target, and if that route runs through a user’s attention or judgment, the frame will take it. This is a property of the design rather than of the model’s intelligence, and it is why user satisfaction makes a weak test: satisfaction is one of the variables a system can optimize.

The fourth is the loss of graceful failure. A tool that is hard to replace is a tool that takes the skill with it when it goes dark. Redundancy, offline operation, and the ability to export one’s own material are dull requirements, and they are what keep the apparatus wearable.

A fifth signature is quieter and has begun to show up in professional work: convergence. When many people lean on the same model for the same task, their drafts, analyses, and even their framings drift toward a common distribution, and the variation that comes from independent struggle narrows. This is plausible engineering supported by suggestive rather than direct evidence, inferred from the behavior of systems tuned on aggregate preference data rather than measured in a controlled trial. The consequence is a collective version of the individual risk. Diversity of approach is itself a form of redundancy, and a team that outsources judgment to one apparatus loses the second opinion that would have caught the apparatus being wrong. An exoskeleton worn by everyone is still worn, but the wearers begin to look alike.

Tests that survive contact with real products

Four questions separate an exoskeleton from a substitute, and none of them requires access to the model’s internals.

Can the person still manage the unaided version, at least roughly? If a writer can no longer draft a page, or a navigator can no longer reach a familiar place, the tool has taken over the task rather than supported it. The relevant measure is not the tool’s output but the person’s capacity when the tool is gone.

Who chooses the objective? A system that serves a user’s stated goals and one that serves a vendor’s metrics can be built from the same model, and the difference surfaces under pressure, when the two goals diverge.

Can the reasoning be inspected at more than one layer? A single confident answer is hard to argue with. A draft, a set of sources, and a sequence of steps can each be questioned, which is what makes correction possible. Tools that expose intermediate work serve a person who intends to stay competent.

Can the person leave? Export, interoperability, and the ability to run a comparable tool elsewhere determine whether the arrangement is a partnership or a dependency. This is a property of the market as much as of the software, and it is the one a user can influence directly by choosing what to adopt.

What an exoskeleton cannot answer

Some of the value in a human life is not a function of capability, and a stronger frame does not manufacture it.

An intelligence can rank options by expected outcomes under a stated objective. It cannot make a contested value judgment stop being contested by being faster at arithmetic. A model can summarize a religious or philosophical tradition, compare its claims, and lay out the arguments on either side, and that is genuinely useful work. It cannot settle which tradition’s account of the good is true, because that question is not the kind that more computation resolves. Capability and legitimacy are different properties, and a system can hold the first in abundance while having nothing to say about the second.

That is why purpose, judgment, and the willingness to answer for a decision remain human work even when the analysis does not. The exoskeleton carries weight. It does not choose the walk.

The line

The metaphor earns its place because it makes a demand that the replacement-brain picture quietly drops. It asks what the person is still doing. An apparatus that leaves someone stronger, better able to see how a conclusion was reached, and free to remove the apparatus passes. A system that decides where to walk, filters what the walker can see, and cannot be taken off fails, however comfortable it feels from the inside.

The power of the technology makes the demand sharper rather than obsolete. The more capable the model, the more consequential it becomes who holds the destination.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home