In 2023, a team of economists ran a controlled experiment with 453 college-educated professionals and gave half of them access to a generative AI assistant. The people with the tool finished writing tasks about 40 percent faster, and independent evaluators rated their output about 18 percent better than the control group’s. The result, published in Science by Shakked Noy and David Zhang, is usually reported as “AI makes workers more productive” (Noy & Zhang, 2023).
The more consequential detail is buried one level down. The gains were largest for the participants who were weakest to begin with, and the tool narrowed the spread between the best and worst performers. The average improvement and the change in the distribution of performance are separate facts. The second one is what the phrase “extraordinary power for ordinary people” actually refers to. It describes a compression of the gap between what a trained professional can do and what an untrained person can do with help.
That compression is measurable, and it has now been measured several times, in several industries, with real workers doing real jobs. What has not been measured is whether the compression makes people powerful in any durable sense. A person who can produce a legal brief, a working script, or a research summary they could not have produced alone has gained access to a capability. Access is not the same as ownership, and a capability that is rented can be withdrawn. This article takes the demonstrated results seriously, uses them to map where the leverage actually lands, and is explicit about where the evidence stops and the speculation begins.
What the experiments actually measured
The strongest evidence comes from randomized trials inside large employers, where workers were assigned to AI-assisted and unassisted groups and the firms could count outputs directly.
At Microsoft, Accenture, and a third Fortune 100 company, a team tracked 4,867 software developers across three trials. Developers with access to GitHub Copilot completed about 26 percent more tasks than the control group, a result the authors report with a standard error wide enough to leave real uncertainty about the exact size (Cui et al., Management Science, 2026). The distributional pattern was sharper than the headline. Junior and newly hired developers completed roughly 27 to 39 percent more tasks. Senior developers gained something closer to 8 to 13 percent. The tool helped most where expertise was thinnest.
A second study followed 5,172 customer support agents at a company that rolled out an AI suggestion system in stages. Average productivity rose 15 percent, measured as issues resolved per hour. Less experienced and lower-skilled agents improved by roughly 30 percent on that measure, with gains up to about 35 percent in the lowest quintile, while the most experienced agents saw small gains in speed and small declines in quality (Brynjolfsson, Li & Raymond, Quarterly Journal of Economics, 2025). The mechanism the authors identified was instructive: the AI was surfacing patterns that the top performers already used unconsciously, which means it was redistributing tacit expertise rather than generating new expertise.
The third study is the one that complicates the story. A team from Harvard, MIT, and Boston Consulting Group gave 758 consultants access to a GPT-4-class model on a set of realistic consulting tasks. On tasks that fell within the model’s competence, the consultants completed 12.2 percent more work, worked 25.1 percent faster, and produced work rated more than 40 percent higher in quality. Then the researchers included one task that looked similar but sat outside the model’s competence. Consultants using AI on that task were 19 percentage points less likely to reach the correct answer than consultants working without it (Dell’Acqua et al., Organization Science, 2026).
That last number is the most important one in the literature, because it describes the shape of the terrain: a “jagged frontier” where AI is dramatically helpful on one side and actively harmful on the other, and where the border is not visible from the inside. The consultants could not tell in advance which kind of task they were doing. Neither can most people.
Capability that used to require a department
The distributional findings explain something that the productivity averages obscure. When a tool raises the floor faster than it raises the ceiling, it transfers a slice of institutional capability to individuals who previously had no access to it.
Consider materials science. In November 2023, Google DeepMind reported a system called GNoME that screened candidate crystal structures at a scale no human team could match, proposing 2.2 million structures below the known stability threshold and flagging roughly 380,000 as promising candidates for synthesis (Merchant et al., Nature, 2023). Screening compounds against a stability model is exactly the kind of work that once required a department: a group with compute, a curated database, and years of patience. In geometry, a related DeepMind system solved 25 of 30 olympiad-level problems on a standard test set, approaching the performance of an average International Mathematical Olympiad gold medalist (Trinh et al., Nature, 2024).
Read those two results for what they demonstrate. They show that pattern-recognition and search over enormous combinatorial spaces, which was institutional work, can now be done at a scale that used to be inaccessible. They do not show that a lone person with a laptop can now produce a new material. GNoME’s output was predictions, and predictions are the cheap part. Cheetham and Seshadri, two materials chemists who examined a sample of the database, found little evidence for compounds that satisfied their three-part test of being credible, novel, and useful, and argued the work should be described as a list of proposed compounds rather than new materials (Cheetham & Seshadri, Chemistry of Materials, 2024).
The distinction matters for anyone trying to reason about individual power honestly. What AI cheapens is the computational component of institutional capability: search, retrieval, pattern matching, first-draft generation, code scaffolding, translation, classification. What it leaves in place is everything that depends on physical infrastructure, accumulated tacit judgment, and legal accountability. A person can now generate ten thousand candidate hypotheses. Someone still has to decide which ones are worth testing, and someone still has to run the experiment, and someone still has to sign their name to the result.
What the tool does not hand over
The experiments are consistent about a second thing: the value of the AI-assisted worker came from their judgment about when and how to use the tool, and that judgment tracked experience. Senior developers gained least because they already knew most of what the model could offer. Consultants who used the model well did so by treating it as a collaborator with a known area of competence, which required them to have a model of its competence. The people who benefited most from AI were, in an important sense, already competent enough to supervise it.
This creates an awkward implication for the phrase “extraordinary power.” If the tool amplifies judgment, and judgment is what the tool does not supply, then the tool amplifies whatever judgment a person brings. Applied to someone with poor judgment, amplification is not obviously a gift. The same system that lets a careful amateur check a diagnosis against the literature lets a reckless one generate a confident, fluent, and wrong answer at scale.
There is a further asymmetry that the lab studies cannot capture. The output of a model is fluent by construction. Fluency is not evidence. A person without the domain knowledge to distinguish a real answer from a plausible-sounding one is not merely a beginner being helped; they are a beginner being handed something that reads like the real thing. That is a genuinely new situation, and it is the reason the quality of verification tools, provenance signals, and citation habits matters more than the raw capability of the models.
Whether the gains stay
The central question for individual power is not whether a person can produce more with the tool, but whether anything remains in the person when the tool is taken away. On that question there is exactly one piece of credible evidence, and it is more encouraging than the rest of the literature.
The customer-support study happened to include a natural experiment. The AI system occasionally failed, leaving agents to work without suggestions for a stretch of time. Workers who had been using the system continued to resolve more issues per hour during those outages than agents who had never had access to it, and the effect was strongest among the workers who had used the AI most and followed its suggestions most closely (Brynjolfsson, Li & Raymond, 2025).
Two things about that result deserve care. First, it is evidence of medium-run persistence, not permanent skill: the outages were interruptions of hours or days, and nobody has shown that the learning survives a year without the tool. Second, the mechanism is not that the AI did the work. Adherence mattered — the agents who learned were the ones who engaged with the suggestions rather than ignoring them, which fits the picture of a system that transmits the habits of high performers to people who have not yet acquired them.
That is the closest thing to a demonstration that AI assistance can transfer capability rather than merely lend it, and it points at a design lesson. A tool that shows its reasoning and invites the user to engage with it is doing something different from one that returns a finished answer. The first can leave a residue of skill. The second, on the available evidence, does not.
The same discount applies to harm
Every capability that becomes cheap for a diligent person becomes cheap for a careless or malicious one. This is not a prediction; it is already observable.
The United States Copyright Office spent two years studying the effects of generative AI on creative industries, received more than ten thousand public comments, and concluded in January 2025 that the technology raises a specific concern about the dilution of human-created work. Its report found that AI systems need no copyright incentive to create, that material generated wholly by AI is not protectable, and that the case had not been made for new legal protection for AI outputs; among its stated worries was that an increase in AI-generated material could crowd out human authorship and undermine the incentive to create (U.S. Copyright Office, Copyright and Artificial Intelligence, Part 2, 2025).
The concern generalizes well beyond copyright. If producing a competent-looking artifact costs almost nothing, then the binding constraint moves from production to attention, and the volume of plausible-looking junk rises. Sophisticated fraud, coordinated spam, and synthetic persuasion all get cheaper on the same curve that makes legitimate assistance cheaper. A society can therefore become more capable and less trustworthy at the same time. Anyone who claims individual empowerment from AI while ignoring the corresponding empowerment of bad actors is telling half the story.
Rented leverage and owned leverage
Here the demonstrated record ends and inference begins, and it is worth marking the line clearly.
What has been demonstrated is that individuals get measurably better at bounded tasks when given model assistance, that the gains are largest for the less expert, and that the benefit reverses when the task falls outside the model’s competence. What has not been demonstrated is that this adds up to durable individual power. The experiments ran for weeks or months, with workers embedded in organizations that supplied the tasks, the standards, and the evaluation. Nobody has shown that a person with model access accumulates capabilities that survive the loss of the model.
That gap matters because of how the capability is delivered. Most people will not hold the weights of the systems they depend on. They will rent access, which means their leverage is conditioned on a contract, a price, a terms-of-service document, and the continued operation of a company they do not control. A person whose legal work, medical second opinion, or business analysis runs through a single provider has acquired a powerful tool and a corresponding dependency. Whether the tool makes them strong or merely well-equipped is a governance question as much as a technical one.
The plausible engineering direction here is unremarkable: portability of data and memory, the availability of capable open models that a person can run and inspect, and interfaces that let someone move their accumulated context between systems. None of that is speculative science. It is ordinary software engineering and business-model choice. Whether it happens is a matter of incentives, and the incentives of a platform company point toward lock-in rather than portability. That divergence, more than any benchmark, will determine whether cheap capability becomes individual power or individual dependency.
A test that survives contact with the technology
It is possible to say something disciplined about all this without pretending to forecast the future. Four questions separate capability that compounds inside a person from capability that is merely being borrowed.
First, does the work leave behind something the person keeps? Finished scripts, corrected reasoning, an inspection trail, a database the user owns — these persist after the subscription lapses. A stream of answers that vanish with the chat window does not.
Second, can the person see why the system said what it said? A recommendation with traceable sources and an explicit uncertainty estimate supports a decision. An authoritative-sounding paragraph with no provenance does not, and the difference shows up precisely when the stakes are high enough that it matters.
Third, who holds the objective? A tool aimed at the user’s stated goal is different from one aimed at engagement, throughput, or compliance, even when the underlying model is identical. The objective is a choice made by someone, and the person using the system is often not that someone.
Fourth, does the person get better at the underlying work? This is the question the labor studies actually bear on. If a junior developer finishes 30 percent more tasks and learns something from doing them, the leverage compounds. If they finish 30 percent more tasks while outsourcing the reasoning entirely, the same number describes a different and worse trajectory. The published results cannot distinguish these yet. Anyone claiming they can is ahead of the evidence.
Where the leverage actually goes
The honest summary of the current evidence is narrower than the enthusiasm and more interesting than the dismissal. AI measurably transfers a slice of institutional capability to individuals, and it does so most strongly for people who have the least of that capability to start with. That is a real and unusual development, because most productivity technologies amplify the people who are already ahead. It also degrades performance precisely where users cannot tell that they have left the tool’s area of competence, which is a hazard with no clean solution.
So the question of what happens when one person can call on a department’s capabilities has an answer that depends on details the studies have not settled. If the capability is portable, inspectable, and connected to real judgment, then ordinary people get something closer to extraordinary power: the ability to attempt work that was previously gated by credentials and budgets, and to keep the results. If the capability is opaque, rented, and aimed at someone else’s objective, then ordinary people get a compelling new interface to dependency, and the gains in measured output will describe a transfer of skill out of people rather than into them.
Both futures are compatible with the numbers currently in the literature. Which one arrives is being decided now, in product decisions about portability, in procurement decisions about single-supplier dependence, and in educational choices about whether to use these tools to do the work or to learn the work. The technology supplies leverage. It does not supply the direction, and it will not tell anyone when it has quietly taken over the steering.
Sources and further reading
- Cui, Demirer, Jaffe, Musolff, Peng & Salz, “The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers,” Management Science, 2026
- Brynjolfsson, Li & Raymond, “Generative AI at Work,” Quarterly Journal of Economics 140(2), 2025
- Dell’Acqua et al., “Navigating the Jagged Technological Frontier,” Organization Science 37(2), 2026
- Noy & Zhang, “Experimental evidence on the productivity effects of generative artificial intelligence,” Science 381(6654), 2023
- Merchant et al., “Scaling deep learning for materials discovery,” Nature 624, 2023
- Cheetham & Seshadri, “Artificial Intelligence Driving Materials Discovery? Perspective on the Article: Scaling Deep Learning for Materials Discovery,” Chemistry of Materials 36(8), 2024
- Trinh et al., “Solving olympiad geometry without human demonstrations,” Nature 625, 2024
- U.S. Copyright Office, Copyright and Artificial Intelligence, Part 2: Copyrightability, January 2025
Loading comments…