In 1984, the educational psychologist Benjamin Bloom published a paper titled “The 2 Sigma Problem.” Its method was unremarkable—a synthesis of existing studies of tutoring—and its result was startling in scale. Compared with conventional group instruction, students taught one-to-one performed about two standard deviations above the class average, roughly the distance between a median student and one in the top few percent. Mastery learning, a cheaper alternative, produced a gain near one standard deviation. Bloom framed the finding as an operational challenge rather than a curiosity: a result that large was a target, because no society could afford a personal tutor for every child.
That paper is worth retrieving now because the obvious objection to mass tutoring—that expert attention is scarce and expensive—is exactly the constraint that large language models appear to relax. If a tireless system can explain photosynthesis at any level of detail, walk a student through a failed proof, and translate between a specific confusion and a standard curriculum, then one of education’s scarcest resources seems to have become abundant. The question in this article’s title stops being rhetorical and becomes empirical: if retrieval is nearly free and explanation is nearly free, what work remains for a learner to do?
Part of the answer is demonstrable in controlled studies. Part is reasonable engineering. Part is open speculation about goals that research cannot settle on its own. Keeping those categories separate is not pedantry, because the practical question—what should schools actually do—depends on which claims are measured, which are extrapolated, and which are choices about values.
Bloom’s two-sigma problem, revisited
Bloom’s magnitude has held up in later synthesis. In a 2011 review in Educational Psychologist, Kurt VanLehn compared the effectiveness of human tutoring, intelligent tutoring systems, and other tutoring approaches. Human tutoring produced an average effect size of about d = 0.79 against no tutoring; intelligent tutoring systems achieved d = 0.76. The gap between a human expert and well-built software was, in that literature, close to zero.
This establishes two things at once. Tutoring genuinely works—its effect is among the largest in education research—and the active ingredient is not obviously the humanness of the tutor. VanLehn’s review suggests much of the benefit comes from structure: presenting tasks at the right difficulty, prompting students to explain their reasoning, and responding to specific errors rather than general confusion. That is a mechanism description, and mechanisms can be implemented in software.
The cautious reading is that Bloom and VanLehn justify engineering optimism, not the conclusion that any chatbot is a tutor. The results attach to systems deliberately built and tested against learning outcomes. Whether a general-purpose language model inherits those gains is a separate, empirical question, and the recent evidence is mixed in a way that turns out to be instructive.
What the best current evidence actually shows
The most rigorous recent test is a randomized controlled trial published in Scientific Reports in 2025 by Gregory Kestin, Kelly Miller, Anna Klales, Timothy Milbourne, and Gregorio Ponti. The setting was a Harvard physics course with 194 undergraduates. The study compared an AI tutor designed to be pedagogically conservative—one that offered hints and guided reasoning instead of answers—against in-class active learning.
The results were substantial. Students in the AI-tutored condition learned more than twice as much in less time; median time on task was about 49 minutes against roughly 60 minutes for the control. Depending on the measure, the reported effect size ranged from about 0.73 to 1.3 standard deviations. This is peer-reviewed, controlled evidence that a carefully constrained model can function as an effective tutor, and it is the strongest such evidence currently available.
It is also narrower than the headlines implied. The trial covered one course at one institution, in a problem-solving subject, using a tutor whose design deliberately suppressed the model’s most fluent behavior—supplying the direct answer. The intervention worked in part because it refused to do the thing that makes chatbots feel most useful.
A second line of evidence cuts the other way and deserves equal weight. In a study of roughly a thousand Turkish high-school students, Hamsa Bastani and colleagues tested two versions of GPT: one unrestricted and one with safeguards steering students toward reasoning. The unrestricted version raised practice performance by about 48 percent, but students who used it scored about 17 percent worse on the later exam than the control group. The safeguarded version nearly tripled practice performance—a gain of about 127 percent—yet still produced no improvement on the exam. The authors described a “crutch” effect: the tool improved performance while present and left learning no better, or worse, once removed. Students also overestimated how much they had learned. The paper appeared in PNAS in 2025.
Set the two studies side by side and the honest summary is conditional. AI tutoring can produce large learning gains when the interaction is designed to make the student perform the cognitive work. It can produce no learning, or negative learning, when it lets the student delegate that work. The decisive variable is not the technology. It is the pedagogy wrapped around it.
A tutor and an answer machine are not the same thing
The contrast between those studies is a concrete version of a distinction the educational literature has drawn for decades: the difference between a tool that supports a cognitive process and one that replaces it. A calculator supports arithmetic after number sense is built; introduced before, it can short-circuit the sense of magnitude that tells you a result is absurd.
The Bastani result is the cleanest demonstration that the concern is not hypothetical. Students had a plausible sense that they were learning—practice scores were high—while the exam revealed that the knowledge had not consolidated. Had education measured only practice performance, the crutch would have been invisible. The exam, which required unaided retrieval, exposed it.
There is a defensible counterargument. If a person will always have the model available in a future workplace, perhaps unaided performance is the wrong target. But that argument assumes availability, accuracy, and the absence of adversarial conditions—assumptions better treated as engineering problems than settled facts. For now, the measured result is that bypassing the learner’s own effort produces measurable deficits, and the size of the deficit depends on how the tool is used.
What offloading actually does to memory
Long before language models, psychologists tested what happens when information is expected to remain retrievable. In a 2011 paper in Science, Betsy Sparrow, Jenny Liu, and Daniel Wegner reported that when people believed a fact would be available later on a computer, they remembered the fact itself less well but remembered where to find it better. The effect—widely called the Google effect—is a demonstration of transactive memory: cognition distributed across people and devices, with the mind tracking sources rather than storing contents.
This is not automatically a harm. Distributed memory is how human groups have always operated; a legal team that remembers which colleague knows the precedent outperforms one that expects every member to memorize the case law. The Sparrow finding shows adaptation, not damage. What it changes is the composition of education. If the mind naturally offloads retrievable facts, then a curriculum built on fact retention trains a capacity the environment no longer rewards.
The stronger claim—that search and AI use degrade thinking generally—has weaker support. A 2025 study in Societies by Michael Gerlich reported negative correlations between AI tool use and critical-thinking scores, but the paper’s data have been publicly questioned for implausible respondent characteristics and inconsistent internal tables, and it should not be treated as established. The defensible position is the measured one from Sparrow: the expectation of retrieval shifts what gets stored. What that shift does to higher-order thinking depends on what education chooses to build, which is a design decision rather than a discovered fact.
Skill, judgment, and the limits of rule-following
If factual recall is being offloaded, the natural candidate for what education should protect is judgment—the capacity to decide what a particular situation calls for. Here the psychology of expertise is instructive, and one model in particular has aged well.
In work beginning with a 1980 Air Force report and developed in the 1986 book Mind Over Machine, Stuart and Hubert Dreyfus described skill acquisition as a progression through stages: novice, advanced beginner, competence, proficiency, expertise. The novice follows context-free rules—shift gears at this speed, exchange pieces by point value. The expert perceives a situation holistically and acts without deliberation. The Dreyfuses argued that the transition happens only through involved experience: the competent performer becomes emotionally invested in outcomes, and that investment is what allows intuition to replace explicit rules. Their controversial corollary was that over-reliance on rules can stall a learner at competence.
The Dreyfus model is a framework rather than a measured effect size, and its strong claims about machine limits have been contested. But its descriptive core is widely used in medicine, nursing, and aviation, and it offers a precise way to say what pure information access cannot supply. An answer machine delivers the content of the rule-following stages. It does not deliver the accumulated, emotionally weighted experience that lets a person recognize that this patient, this contract, or this student is the exception to the rule. The knowledge an expert draws on when acting well is partly tacit; Michael Polanyi’s observation, cited throughout the expertise literature, is that we can know more than we can tell.
For education, the implication is that the value of practice is not the residue of facts it leaves behind. It is the calibration of judgment—learning which intuitions to trust, which to distrust, and when a confident answer is likely to be wrong. That calibration is slow, and it is exactly what an always-available answer can tempt a learner to skip.
What education is for
UNESCO’s 2023 Guidance for generative AI in education and research, authored by Fengchun Miao and Wayne Holmes, reaches a similar structural conclusion from a policy direction. The document argues that generative AI “replicate[s] the higher-order thinking that constitutes the foundation of human learning,” and that this forces institutions to revisit “why, what and how we learn.” Its concrete recommendations are modest—data protection, an age threshold of 13 for unsupervised chatbot use, validation of tools for pedagogical appropriateness—but its conceptual move is the important one. Once finished answers are abundant, the center of gravity of education shifts toward the capacities that let a person evaluate, frame, and choose.
Those capacities can be named. The first is error detection: recognizing when a fluent answer is wrong, biased, irrelevant, or subtly off requires enough domain knowledge to notice the mismatch, which is why “just ask the AI to check its own work” is circular. The second is problem framing: an external intelligence answers the question it is asked, and forming the right question is a skill that cannot be outsourced to the thing being asked. The third is transfer: applying an idea in a new context, which is precisely what the exam in the Bastani study measured and the practice scores did not. The fourth is judgment about values: deciding which goal matters, which is not a retrieval problem at all.
None of these is a new educational aim. What is new is that the older aim—broad factual coverage—has lost its monopoly, because the cost of retrieving a fact has collapsed. A curriculum that still spends most of its time on retrieval is now spending it on the cheapest thing available. The shift is not from knowledge to ignorance; it is from knowledge as the possession of answers to knowledge as the capacity to use, test, and question them.
The opportunity, taken seriously
The caution in the preceding sections should not obscure the genuine upside, which is large. For most of educational history, the quality of instruction a student received depended on where they happened to live and which teacher they happened to draw. Bloom’s one-to-one advantage was real but rationed by cost. A system that can supply patient, available, level-appropriate explanation to any student with a device compresses that inequality of access in a way no previous technology did.
The gains are not confined to remediation. When the mechanical parts of a subject can be handled by a tool, class time can move toward the parts that require discussion, disagreement, and construction of an argument—activities that are pedagogically valuable precisely because they are hard to automate. A teacher freed from delivering the same explanation forty times can spend the hour on the specific misconception in front of them. The Kestin trial is evidence that this division of labor can work when the tool is designed for it.
There is also a defensible case that some skills deserve to be offloaded. If a surgeon will always have a reference system available, requiring memorization of every drug interaction may be a poor use of a decade of training. The question is not whether to offload but which capacities must remain load-bearing in the person—because those are the ones on which the person’s judgment depends when the tool is unavailable, wrong, or hostile.
The engineering answer, then, is the same as the pedagogical one: put the tool where it supports the process and remove it where it replaces it. That is a design constraint, not a slogan, and it requires measuring the right outcome. The Turkish study showed how easily that measurement goes wrong.
Assessment is where the argument becomes concrete
If the purpose of education shifts from coverage to judgment, the machinery of assessment has to shift with it, and this is where abstract agreement often breaks down. A multiple-choice test of facts is easy to administer and easy to game once a model can answer it. An essay is easier to generate than to falsify. The resulting temptation is to ban the tool, which removes the benefit along with the cheating, or to permit it, which removes the evidence that learning happened.
A more durable option treats the tool as part of the exam. If students may use a model, the assessment can require them to evaluate its output—to identify the error it introduced, to defend a claim it made, to detect the planted bias. This tests the capacities that remain scarce precisely because the tool is present, and it is closer to the conditions graduates will actually face. It is also harder to design and harder to score, which is why institutions under cost pressure may avoid it.
The Bastani exam result is the reason this section matters. The practice scores looked excellent; the unaided exam revealed the gap. Any assessment that cannot distinguish performance-with-support from performance-without it will systematically overstate what students have learned. That is a measurement problem before it is a moral one, and it is fixable—but only if someone is willing to define what unaided competence should look like in a world where unaided competence is no longer automatic.
What the evidence does not settle
The limits of the evidence should be stated plainly. The Kestin trial covers one subject at one university. The Bastani study covers one national school system and one generation of models. Bloom’s synthesis is four decades old. The UNESCO guidance is a normative document, not a controlled result. From these it does not follow that AI tutors will or will not transform education; it follows that their effect is mediated by design choices still being made.
What the evidence does not support is the collapse of education into access. If the argument were “the model knows more, therefore learning is obsolete,” the two studies that measured learning would have to be irrelevant, and they are not. A student who can generate an answer has not thereby acquired the ability to judge it, and a student who cannot judge it cannot detect the moment the model is confidently wrong. Reliance becomes riskier as the answers improve, because improvement makes deference feel more reasonable.
Bloom’s two-sigma gap was a problem of scarce teachers. The new gap is different: it is the distance between access and understanding, and it widens when the tools make access feel like understanding. If that gap is closed at all, it will be by the unglamorous work of designing interactions that leave the cognitive effort on the learner’s side. The technology sets the cost of the answer. It does not determine what the answer is worth to someone who cannot evaluate it.
Sources and further reading
- Bloom, B. S. (1984). “The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring.” Educational Researcher, 13(6), 4–16. https://doi.org/10.3102/0013189X013006004
- VanLehn, K. (2011). “The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems.” Educational Psychologist, 46(4), 197–221. https://doi.org/10.1080/00461520.2011.611369
- Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports, 15, 17458. https://www.nature.com/articles/s41598-025-97652-6
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). “Generative AI Can Harm Learning.” Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.2422633122
- Sparrow, B., Liu, J., & Wegner, D. M. (2011). “Google Effects on Memory: Cognitive Consequences of Having Information at Our Fingertips.” Science, 333(6043), 776–778. https://doi.org/10.1126/science.1207745
- Dreyfus, S. E. (2004). “The Five-Stage Model of Adult Skill Acquisition.” Bulletin of Science, Technology & Society, 24(3), 177–181. https://doi.org/10.1177/0270467604264992
- Miao, F., & Holmes, W. (2023). Guidance for generative AI in education and research. UNESCO. https://doi.org/10.54675/EWZM9535
Loading comments…