The chat window answers in under three seconds. A student stuck on a problem set asks why a confidence interval widens when a sample shrinks, and a clean explanation appears: the standard error scales with one over the square root of the sample size, so less data means more uncertainty about the estimate. The student reads it, feels the click of comprehension, and moves to the next problem. The session felt efficient. Almost nothing in it was hard.
Compare that with the same student working the reasoning alone: sketching the sampling distribution, guessing how the width changes, being wrong in a specific way, and only then checking a text to find out why. The second session is slower and more irritating, and it produces a memory that lasts. The first is smooth and pleasant and frequently leaves nothing behind.
Two activities that look identical
From the outside, comprehension and capability are nearly indistinguishable. A learner who reconstructs an argument from memory and a learner who reads the same argument off a screen produce the same sentences and the same nod of understanding. The difference only becomes visible later, when the screen is gone.
Cognitive scientists describe this as the gap between fluency and retrieval. Fluency is the ease of processing information while it is in front of you. Retrieval is the ability to reconstruct it when it is not. An assistant that explains clearly trains fluency. An assistant that withholds until you have wrestled with the problem trains retrieval. The same model can do either, which means the choice between them is a design decision rather than a property of the technology.
That reframing matters because the obvious reading of a powerful assistant is that it makes learning faster. In some configurations it does. In others it makes the feeling of learning faster while degrading the thing being felt.
What the classroom experiments actually found
The clearest evidence about AI and durable learning comes from a randomized controlled trial in Turkish high schools, published in PNAS in 2025. Bastani and colleagues gave roughly a thousand students access to a GPT-based tutor during mathematics practice. One group used an unrestricted assistant that would supply answers. A second used a version engineered to guide with hints and prompts rather than solutions. A third worked unaided.
During practice, the unrestricted group outperformed the control by 48 percent, and the guided group by 127 percent. Then the researchers took the AI away and tested what the students had retained. On the exam, the unrestricted-assistant group scored 17 percent below the students who had never used it at all. The guided group performed about the same as the control. The tool that made students fastest during practice produced the worst durable learning, and it did so without the students noticing: they believed they had learned more than the control group. They were not careless. They were reading the fluency signal and misinterpreting it, which is precisely what the fluency illusion predicts.
A second study approaches the same mechanism from the opposite direction. Kestin and colleagues, also in 2025 in Scientific Reports, built an AI tutor around established learning principles and tested it in a Harvard physics course. Students who used it achieved median learning gains more than double those of students in an active-learning classroom, with effect sizes above one standard deviation on some measures. The tutor did not simply answer questions. It was designed to elicit predictions and explanations first, forcing students to commit before receiving feedback. The pairing is instructive: withholding the answer helped, and handing it over hurt.
Both results deserve to be held firmly and lightly at once. Each is a single trial in a single subject over a short horizon, measuring test performance rather than a career of thinking. But the internal logic is unusually clean. The variable that changed was how much of the answer the student had to generate, and the outcome moved in the opposite direction from the feeling.
Why effort is the mechanism
The pattern predates large language models by decades. In a 2006 experiment, Roediger and Karpicke had students read passages and then either restudy them or take a practice test. Immediately afterward, the restudiers felt more confident and scored higher. A week later the order reversed: the tested group recalled roughly 61 percent of the material against about 40 percent for the restudiers, and testing helped even when no feedback was given. Restudying inflated confidence without producing durable memory.
A 2013 monograph by Dunlosky and colleagues reviewed ten widely used study techniques and rated their usefulness. Two earned the highest marks: practice testing and distributed practice, both of which impose effortful retrieval and spacing. Among the techniques rated least useful were rereading and highlighting, which feel productive and largely are not.
A related result concerns the feeling of learning itself. Deslauriers and colleagues, in a 2019 PNAS study, compared identical content delivered as polished lectures or as active learning. Students learned more in the active condition but reported that they had learned less, and they rated the lectures more highly. The discomfort of effort was being read as evidence of poor teaching. The authors called this the fluency illusion, and it is the same misreading that led the Turkish students to trust the unrestricted tutor.
The underlying principle is what Robert Bjork and Elizabeth Bjork named desirable difficulty: conditions that slow performance during acquisition while improving long-term retention and transfer. Retrieval practice, spacing, interleaving, and varied practice all qualify. They share a signature — they make the present harder and the future better.
When difficulty stops helping
It would be a mistake to conclude that harder is always better. Difficulty is desirable when the learner has the background to struggle productively and the feedback to correct course. Below that threshold it becomes mere obstruction.
For genuine novices, guidance often beats discovery. Worked examples — studying a completed solution step by step before attempting your own — can outperform unguided problem-solving early on, because the learner cannot yet generate the right moves and so cannot learn from flailing. Cognitive load theory makes the same point from the other side: a task that exceeds working memory teaches frustration rather than skill.
There is a further wrinkle, the expertise-reversal effect. Instructional support that helps a beginner can actively hinder someone more advanced, who no longer needs the scaffold and is slowed by it. The implication for an assistant is that the correct amount of withholding is a moving target. A system that always withholds is as badly calibrated as one that always answers. Calibration, not maximal friction, is the goal.
This is why generic advice about “productive struggle” is thin. The useful question is empirical and personal: at this level of skill, on this task, does removing the hint improve what I can do a week later? Sometimes the honest answer is that the friction is teaching nothing and should be skipped.
The assistant as a spotter
A spotter in weight training does not lift the bar. The spotter makes heavier lifting safe enough to attempt, and steps in only when the attempt is failing. The plausible engineering for a learning assistant follows the same logic — not by limiting capability, but by structuring when capability is deployed.
Generation before explanation comes first. The learner commits to a prediction, a derivation, or a summary before the model offers its own. Whatever follows is then compared against a produced artifact rather than absorbed passively. This single change is the clearest dividing line between the Kestin tutor and the unrestricted one, and it has the strongest support in the evidence.
Second is Socratic withholding. The assistant responds to a request for an answer with a question about the reasoning that led there, escalating to a full explanation only after a real attempt. This should not be a universal rule; sometimes the fact is trivial and the friction is pure waste. Knowing when to withhold is part of the skill of using the tool, and it can be tuned.
Third is scheduled retrieval. An assistant that answers every question on demand can never ask about yesterday’s topic. One that revisits earlier material at expanding intervals, in a different context and without warning, supplies the spacing and interleaving that the research consistently rewards.
Fourth is error harvesting. The most informative moments are the ones where the learner was wrong for a reason. An assistant that records the specific misconception, not just the corrected answer, builds a map of where understanding is thin. That map is far more useful than a completion percentage, and it is something a person can act on.
Dependence and the erosion of skill
There is a second failure mode, separate from the fluency illusion. It is the quiet loss of a capability the learner once had. Aviation offers the sharpest documented case.
Modern aircraft are flown mostly by automated systems, which improved safety and precision. By 2013 the Federal Aviation Administration had concluded that continuous reliance on those systems was degrading pilots’ ability to recover an aircraft when the automation handed control back at the worst possible moment. That January the agency issued a safety alert encouraging operators to build deliberate opportunities for manual flight into routine operations — because the autopilot never asked pilots to practice the skill the autopilot had replaced.
The analogy is structural rather than exact. A pilot’s manual control is a discrete, trainable skill with a visible failure signature, while judgment and reasoning are diffuse and hard to measure. But the causal shape is the same: a system that performs a task competently removes the occasions on which the human would otherwise exercise it. If those occasions were the training, the training quietly stops, and no one notices until the situation demands the missing skill.
The honest counterargument is that offloading is frequently correct. Nobody insists a modern statistician compute square roots by hand, and delegating a mechanical step to a reliable tool is itself a competence. The test is whether the delegated step is load-bearing for the skill you care about. Offloading arithmetic frees attention for modeling. Offloading the modeling removes the thing you were trying to build.
Incentives and whose objective
There is a design pressure that runs against everything above. Answer-first interaction is what users reward. It feels like the tool is working, it produces immediate satisfaction, and satisfaction drives retention, subscription, and engagement. A system optimized for how helpful it feels will drift toward supplying answers, because that is the behavior that gets loved.
This is not a conspiracy; it is an incentive gradient. The same dynamic shapes tutoring software, search, and every other interface measured by session length and return visits. Fluency is easy to sell because it produces a pleasant feeling in the moment, and the costs land weeks later, off-platform, invisible to the metric that governs the design.
A person can partly defend against this by choosing the interaction deliberately, asking for critique instead of completion, and treating the smoothest sessions with suspicion. But the defense is imperfect. The absence of friction is the signal that the effort — and therefore the learning — has been moved elsewhere.
Where the evidence thins out
The evidence for retrieval practice, spacing, and the fluency illusion is strong and old. The evidence for AI specifically is young, thin, and concentrated in short educational trials. Almost no published study follows learners for months, measures transfer to unfamiliar domains, or tracks what happens when the assistant is removed for good. The Kestin result is encouraging and describes one carefully engineered system in one course. The Bastani result is sobering and describes one subject across one term. Neither settles the general case, and it would be overreach to treat either as final.
Transfer is where the ground gets softest. Barnett and Ceci, reviewing a century of transfer research in 2002, argued that asking whether transfer happens is the wrong question; transfer varies continuously along dimensions such as how far knowledge must travel and how different the new context is. Near transfer, to a task closely resembling the training, is common. Far transfer, to a structurally different domain, is rare and heavily dependent on conditions that are often untested. A tutoring system that reliably improves performance on its own problem set has demonstrated near transfer at best. Whether it builds judgment that travels is, for now, a plausible hypothesis rather than a measured result.
What would count as progress
Because fluency misleads, the measures that matter are the awkward ones. Can the learner solve a new problem without hints several days later, in a context that looks nothing like the practice? Can the learner catch the model’s error rather than accept it? Can the learner explain the idea to someone else with the tool closed?
An honest account has to concede that these questions are hard to instrument and easy to fake. A system can be built to report gains that flatter its own design, and self-reported confidence is exactly the signal that should be distrusted. The safeguard is a willingness to test the learner with the assistant switched off — the one condition a well-intentioned assistant has an incentive to avoid.
The larger question
The live frontier is calibration. A person with a powerful assistant and preserved skill can do work neither could do alone. A person with a powerful assistant and atrophied skill has outsourced the part of themselves that would have known whether the output was any good. The first relationship compounds. The second merely feels like it does.
Getting there means keeping the load where it belongs: on the person for the operations that constitute the skill, on the machine for the ones that do not. That boundary is not fixed, and much of it remains an open empirical question. But it is testable, and the tests already exist. The only thing required is a willingness to run them with the tool turned off.
Sources and further reading
- Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. “Generative AI without guardrails can harm learning: Evidence from high school mathematics.” Proceedings of the National Academy of Sciences, 2025. https://www.pnas.org/doi/10.1073/pnas.2422633122
- Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. “AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting.” Scientific Reports, 2025. https://www.nature.com/articles/s41598-025-97652-6
- Roediger, H. L., & Karpicke, J. D. “Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention.” Psychological Science, 2006. https://journals.sagepub.com/doi/10.1111/j.1467-9280.2006.01693.x
- Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. “Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom.” Proceedings of the National Academy of Sciences, 2019. https://www.pnas.org/doi/10.1073/pnas.1821936116
- Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. “Improving Students’ Learning With Effective Learning Techniques.” Psychological Science in the Public Interest, 2013. https://journals.sagepub.com/doi/10.1177/1529100612453266
- Barnett, S. M., & Ceci, S. J. “When and where do we apply what we learn? A taxonomy for far transfer.” Psychological Bulletin, 2002. https://pubmed.ncbi.nlm.nih.gov/12081085/
- Federal Aviation Administration. “SAFO 13002: Manual Flight Operations.” 2013. https://www.faa.gov/sites/faa.gov/files/other_visit/aviation_industry/airline_operators/airline_safety/SAFO13002.pdf
Loading comments…