AI · Article 56 of 64

Will AI Make Us Smarter—or Make Us Helpless?

Can human-plus-AI capability rise while unaided human capability falls?

Anyone who has driven with a navigation app for a decade and then tried to cross an unfamiliar city without one knows the small shock of it. The route is not especially difficult, yet the mental map that once assembled itself so easily refuses to appear. A capacity that felt permanent has gone quiet. The discomfort is rarely about geography alone. It is the discovery that something you assumed was part of you had been quietly delegated, and that the delegation arrived with a slow settling of the thing itself.

That moment is the everyday version of a question this series keeps circling. What a person can accomplish with a tool in hand and what they can accomplish alone are two separate quantities, and nothing guarantees they move together. Someone can become more productive, more fluent in unfamiliar fields, and more willing to attempt hard problems while the set of things they can do unaided contracts. Whole populations can do the same. The question worth taking seriously is whether AI widens the distance between those two curves more than earlier tools did, and whether the widening matters.

A qualified answer is the honest one. Nothing about AI forces unaided ability downward, and a thoughtfully built tool can raise it. But the incentives around using these systems push in one direction, and the evidence from navigation, from industrial and aviation automation, and from the first large studies of generative AI at work all lean the same way. Augmented capability rises quickly and visibly. Unaided capability answers to practice, which is slow, costly, and rarely measured. The distance between the two stays hidden as long as the tool is working.

Two quantities measured on very different clocks

The first quantity is what a person can accomplish with the system in hand. Call it augmented capability. It is easy to observe: output per hour, problems solved, drafts produced, hypotheses tested. Employers, markets, and research labs all have reasons to track it, and its gains show up in revenue and publication long before any careful study is done.

The second quantity is what the same person can do when the system is gone. Call it unaided capability. Measuring it means deliberately withholding the tool and testing the person, which is expensive, unpopular, and pointless by the standards of next quarter’s results. So it gets measured rarely, and usually by accident — when a pilot has to hand-fly, a clinician has to read a study alone, or a driver’s phone dies on the far side of town.

A third quantity matters even more and is measured least of all: the stock of skill held by a population over time. A generation can keep operating a system at a high level while the underlying competence that built it drains away, because the people who hold that competence are still present. The problem arrives when they leave and their replacements learned the work only in its automated form. Lisanne Bainbridge noticed this in 1983 and treated it as one of automation’s central ironies rather than a footnote.

What the navigation studies actually show

The navigation case is the most vivid, and it is worth being precise about how strong the evidence is.

In a 2020 paper in Scientific Reports, Louisa Dahmani and Véronique Bohbot surveyed 50 participants about their lifetime use of GPS-based navigation and their preferred way of finding their way — by landmarks, by a mental survey of the environment, or by following turn-by-turn instructions. The participants then completed a battery of spatial memory tasks, including measures known to depend on the hippocampus. Greater reported lifetime GPS use was associated with worse performance on the hippocampal-dependent measures after accounting for age, gender, video-game experience, and other factors. In a subset of 13 participants who returned roughly three years later, those who reported more GPS use over the interval showed a steeper decline.

The authors argue the direction of causation runs from habitual GPS use to weaker spatial memory, and they offer a mechanism: relying on turn-by-turn guidance removes the occasions on which a person would practise building a cognitive map, and the brain keeps what it uses. The counter-explanation is equally plausible on its face, since people with a weaker sense of direction may simply reach for GPS more often. The authors used the longitudinal subsample to argue against it.

The honest reading is that this study is consistent with a deskilling story and does not prove it. Fifty people is a small sample, the design is largely cross-sectional, and lifetime GPS use is self-reported. A single laboratory’s finding does not establish a civilisational trend. What it does is make the mechanism concrete enough to test, and it removes the comfortable assumption that spatial memory is immune to how we move through the world.

Remembering less, or remembering differently

Four years before that, in Science, Betsy Sparrow, Jenny Liu, and Daniel Wegner ran four experiments on what they called the Google effect. Participants who faced difficult trivia questions were primed to think about computers and search engines. When new information was presented as something they would be able to look up later, they remembered the content less well and remembered where to find it better. When they were told the information would be erased, they remembered the content more strongly. Their memory appeared to be allocated toward whatever would prove useful to retrieve.

That result is often summarised as evidence that search engines are making us forget. It says something narrower. Retrieval expectations shape encoding, which is a long-established principle of memory and the foundation of what psychologists call transactive memory — knowing which colleague knows what, or which book holds which fact. Humans have offloaded to partners, notes, and libraries for as long as they have had partners, notes, and libraries, and this is often a strength rather than a failure. Mary Potter of MIT described the Sparrow results at the time as suggestive rather than conclusive, and that judgment has held up.

The interesting worry sits one level down. If people remember locations instead of content, that is a shift in allocation, and it may cost nothing at all. If people stop practising the reasoning that recall and reconstruction support — the inference, the comparison, the reconstruction of a chain of thought — then something more consequential is happening than a change of filing system. Distinguishing a harmless reallocation from a real loss requires testing unaided performance over time, which almost nobody does.

A 1983 warning that reads like a memo from this year

The clearest framework for thinking about all of this was written before most readers of this article were born. In a five-page paper in Automatica, Lisanne Bainbridge laid out what she called the ironies of automation.

The first irony concerns which tasks remain. Automation is attractive because designers want to hand the simple, reliable, easily specified parts of a job to machines. What is left to the human, by construction, is the residue the designer could not specify: the rare, ambiguous, high-consequence decisions. Removing the easy work therefore makes the remaining work harder rather than easier.

The second irony concerns what that does to the person. The skills needed to handle the residue are maintained by using them, and automation is precisely what removes the routine occasions for use. Physical control skills fade. Attentional habits go stale. The operator is asked to intervene in exactly the situations for which they have had the least recent practice.

Bainbridge also noted that sustained monitoring is a task humans perform badly, with vigilance degrading after roughly half an hour of watching a display on which little happens, and that the operators of any given era are partly riding on the skills of the people who did the work before it was automated. Barry Strauch’s 2018 retrospective in IEEE Transactions on Human-Machine Systems traced three decades of accidents in aviation, marine operations, and process control back to those unresolved ironies. It is a useful corrective to the assumption that the problems of automation were discovered recently, or that better interfaces will dissolve them.

Hands and heads do not decay at the same rate

If the worry were simply that disused skills fade, the story would be straightforward. The aviation research shows it is more selective than that, and the selectivity is the interesting part.

In a study published in Human Factors in 2014, Stephen Casner and colleagues put 16 experienced airline pilots through routine and non-routine scenarios in a Boeing 747-400 simulator while varying how much automation they used. Instrument scanning and manual control — the hand-eye skills — held up reasonably well, even among pilots who reported practising them infrequently. Cognitive tasks required for manual flight fared worse. When pilots had to track the aircraft’s position without a map display, decide which navigational steps came next, or recognise an instrument system failure without automated cues, problems appeared often and were sometimes large. Pilots who reported more task-unrelated thought while the automation was flying performed worse on those cognitive tasks.

Two implications follow. Skills learned to a high level and kept in some kind of use are durable. What erodes are the skills that require a person to be mentally present, and mental presence is a behaviour rather than a property of the equipment. A crew can be physically in the loop and cognitively out of it, and the automation will not know the difference.

What knowledge workers report about their own thinking

The aviation evidence is decades old and comes from a tightly regulated profession with excellent failure data. The equivalent evidence for generative AI at work is young, and it arrives mostly as self-report.

In a study presented at CHI in 2025, Hao-Ping Lee and colleagues surveyed 319 knowledge workers about 936 concrete examples of using generative AI on the job. Two patterns stood out. Participants who reported higher confidence in the AI’s abilities also reported engaging in less critical thinking; participants with higher confidence in their own expertise reported engaging in more. The kind of thinking that remained shifted from generating answers toward verifying, integrating, and taking responsibility for the output.

This is a survey, so it describes what people believe about their own mental effort rather than what a test of unaided skill would show. Its value is that it locates where the effort moves. If the work of thinking migrates from producing a first answer to checking someone else’s, then the skill that matters most becomes verification — and verification is only meaningful when the person doing it can independently judge the answer. A verifier who has never built the thing being verified is checking against confidence rather than against evidence.

Being in the loop is a behaviour, not a setting

Several lines of research converge on the point that nominally sharing control with automation is not the same as exercising judgment over it.

Pnina Gershon and colleagues analysed a month of real-world driving with a Cadillac Super Cruise system, across 5,514 miles and 265 automation-initiated disengagements. Immediately after the system handed control back, drivers spent less time looking at the road and more time looking at the instrument cluster, and their on-road gaze durations fell sharply at the upper end of the distribution. When drivers had both hands off the wheel, their takeover was slower than when at least one hand was on it. It took several seconds on average before their glance behaviour began to recover. The system had worked as designed, and the driver’s readiness had drifted anyway.

A related caution comes from Yueying Chu and Peng Liu’s 2023 review in Ergonomics, which argues that “automation complacency” is frequently invoked to blame drivers in crash investigations on thinner scientific ground than the term’s confident use suggests. Their point is not that people never over-trust automation. It is that a concept used to assign blame should be measured carefully, and that design choices and expectations bear a share of the responsibility. Both conclusions transfer to the AI debate, where “the user over-relied on the model” can serve as an all-purpose excuse for systems that were hard to monitor in the first place.

The asymmetry that decides how bad this gets

Whether the gap between the two curves becomes dangerous depends heavily on who is affected.

A senior professional who built a skill by hand before the tools arrived retains a foundation to fall back on. They can partly reconstruct what the model did, recognise when an answer is shaped wrong, and absorb the loss of practice for a while. A junior who has only ever worked with the tool has no such reserve. Bainbridge’s generational irony applies directly: the competence that keeps a system safe is often embodied in the people who are closest to retiring.

The paradox of verification sharpens this. The CHI survey suggests that as routine generation is delegated, the residual human work becomes checking and integrating. But checking depends on the competence that doing the work built. Organisations that route every first pass to a model while telling themselves that humans still supervise are, over time, thinning out the population of people capable of supervising. The supervision becomes formal and the judgment becomes nominal.

Sorting what is demonstrated from what is plausible

The evidence is strong enough to support some claims and not others, and it pays to keep them apart.

Demonstrated, with the caveats noted above: human vigilance on rare-event monitoring decays within tens of minutes; unused skills decay, with cognitive and procedural skills decaying faster than well-learned motor skills; heavier reported GPS use is associated with weaker hippocampal-dependent spatial memory; pilots lose the cognitive skills of manual flight faster than the hand-eye skills; drivers in partially automated cars take measurably longer to re-engage and need several seconds to recover their scanning; and knowledge workers report spending less deliberate thinking when they trust the model more.

Plausible engineering, not yet demonstrated at scale: interfaces that make unaided practice routine, systems that withhold answers in order to build a user’s capacity, difficulty that adapts to keep a person at the edge of their competence, and verification tools that supply independent evidence rather than reassurance. These are design proposals with reasonable mechanisms behind them, and the evidence for and against them is still being gathered.

Speculation: the claim that AI is producing a permanent, civilisation-wide decline in human intelligence. Nothing in the current literature supports a trend of that size, and the instruments needed to detect one — longitudinal, unaided, population-level measures of specific competencies — are mostly not being collected.

Keeping both curves rising

The gap between augmented and unaided capability is not a law of nature. It closes or widens according to choices that remain available.

The first is to treat practice as part of the job rather than an obstacle to it. Aviation did this, imperfectly and late, by writing manual flying back into training and line operations once the evidence accumulated. Professions that adopt AI without an equivalent policy are accepting the same drift on a faster timeline.

The second is to make verification real. A verification step is only as good as the evidence available to the verifier, which means systems should be expected to expose sources, assumptions, and uncertainty rather than polished conclusions. Verification that consists of asking the model whether it is sure trains nothing and catches little.

The third is to distrust fluency as a measure of competence. The most seductive way a capable tool fails is by making its users feel more competent than they are, which suppresses the very discomfort that would prompt them to practise. The navigation study, the pilot study, and the knowledge-worker survey describe different versions of that mismatch.

The fourth is to design for the moment the tool is absent. Systems that fail gracefully, that leave behind a person with more options than they started with, and that can be inspected rather than merely trusted are the ones that keep augmented and unaided capability from diverging. That is a design criterion with consequences: for how models explain themselves, how much they cultivate dependence, and how much room they leave for practice.

Where this leaves the question

The two curves can separate, and there is enough evidence to say that they have been separating in specific, measurable places for a while. Nothing guarantees the separation becomes catastrophic, and nothing about it is inevitable in the strong sense. What the record shows is that the drift is quiet, that the incentives run toward encouraging it, and that it stays invisible in exactly the metrics organisations prefer to track.

The useful question for a person or an institution is narrower than whether AI makes people smarter in general. It is which specific unaided capacities you are counting on, whether anything in your daily work keeps them exercised, and what would happen on the day the tool is unavailable, wrong, or simply not trusted by the person who has to sign off. Those are answerable questions, and answering them is what keeps a system a tool rather than a substitute for the person using it.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home