AI · Article 22 of 64

Can We Use AI Today to Become Smarter?

Without brain implants, can current AI meaningfully increase human capability?

Without brain implants, can current AI meaningfully increase human capability?

In 2023, a large introductory physics course at Harvard ran a randomized comparison whose results are now cited far more confidently than they were originally reported. Greg Kestin and his colleagues assigned students to two versions of the same lesson. One group learned through the course’s usual in-class active-learning exercises; the other worked with an AI tutor built specifically for the material. The students who used the tutor posted median learning gains more than twice as large as the control group, and they finished in less time — roughly forty-nine minutes of work against sixty. When the study reached Scientific Reports in 2025, it became one of the strongest demonstrations yet that ordinary software, with no electrode near a skull, can measurably change how fast a person learns something real.

That result is easy to overstate. Turning it into “AI makes you smarter” collapses several different questions into one. Whether a system can outperform a worksheet at a delimited task is not the same as whether its gains accumulate inside the person: whether they survive the screen closing, whether they transfer to problems the system never saw, and whether the learner ends up more capable rather than merely better served. The honest answer to the question in this article’s title is that current AI can raise a person’s effective capability by a meaningful margin, and it can also lower it, depending largely on how the tool is aimed. The difference between those outcomes is not mysterious. It has already been measured.

A tutor that never gets tired

The Harvard physics experiment is worth unpacking because its design tells us where the effect came from. This was not a chatbot answering homework questions. It was a tutor configured to prompt, correct, and sequence problems, embedded in a course where the comparison group was already receiving good instruction rather than a passive lecture. The control condition was active learning, the sort of teaching that most education research treats as a strong baseline. Beating it by a factor of two is a bigger claim than beating a video.

The mechanism is not exotic. Decades of research on human tutoring have identified the ingredients that matter: immediate feedback, forced retrieval rather than passive review, the pacing of difficulty to the edge of the student’s current ability, and enough interactivity to keep attention on the task. A skilled tutor delivers these adaptively, noticing that a wrong answer reflects a misconception in one student and a slip in another. That adaptability is expensive to provide at scale with humans. It is comparatively cheap with a language model, and the Kestin study is direct evidence that the cheap version can reproduce a substantial part of the benefit in at least one domain.

Note what the study does not establish. It does not show that AI tutoring improves long-term retention in general, that it works equally well across subjects, or that it substitutes for a teacher. It shows a large, well-measured gain on physics learning in a specific, well-designed intervention. Treating that as a general law about cognition is the first place the popular interpretation goes wrong.

When the same tool teaches the wrong lesson

A 2025 field experiment in Turkey gives the mirror image. Hamsa Bastani and colleagues worked with roughly a thousand high-school students in mathematics and tested two ways of using a GPT-based assistant. The first, called GPT Base, was a standard chatbot that would simply help with the problems. The second, GPT Tutor, was designed by the researchers and the teachers to give hints and scaffolding without handing over answers — closer, in spirit, to the Harvard tutor.

The results split sharply. Students using GPT Base became substantially better at the practice problems, improving their scores by a large margin during the sessions. On the unassisted exams that followed, however, they performed about seventeen percent worse than students who had not used the tool at all. The benefit vanished precisely when the machine was taken away, and it left them below where they started. The GPT Tutor group showed no such penalty, and in fact improved their practice performance even more.

This is the central practical finding in the whole literature on using AI to get better at thinking. The same underlying model produced opposite outcomes depending on whether it was allowed to do the cognitive work or was forced to elicit it. A tool that supplies the answer short-circuits the retrieval and struggle from which learning is actually built. A tool that withholds the answer while helping the learner get there can amplify the process. The capability does not live in the model. It lives in the interaction, and the interaction can be designed badly.

The Turkish experiment also exposes a trap in how we measure these tools. By the metric of the practice session — the score on the problems the students were doing with the model — the unguarded chatbot looked like a clear win. Only the unassisted exam revealed that the gain was an illusion. Any evaluation that stops at the moment of use will rank the harmful design above the helpful one. That is a general lesson about AI and capability: the metric that is easiest to collect is usually the one that flatters the tool. Hours saved, tasks completed, words generated — all of these rise whether or not anything is being learned. The measures that actually track capability are slower and more annoying to gather, which is exactly why they get skipped.

A related result is worth adding here. A randomized experiment by Aidan Toner-Rodgers on materials scientists found large productivity gains from an AI assistant in idea generation, but the gains were concentrated among the top researchers, with the bottom half seeing little benefit — a reminder that “AI raises performance” is often a statement about who was already poised to use it well.

The jagged frontier

The most useful single framing of AI’s effect on skilled work comes from a preregistered field experiment by Fabrizio Dell’Acqua and colleagues, run with 758 consultants at Boston Consulting Group. The consultants tackled a set of tasks, some of which fell inside what the researchers called the “jagged frontier” — the shifting, irregular boundary of what the model could do well — and some of which fell deliberately outside it.

Inside the frontier, the results were dramatic. Consultants using AI completed 12.2 percent more tasks and worked 25.1 percent faster, and their solutions were rated significantly higher in quality. The system behaved like a genuine capability multiplier. Outside the frontier, the sign flipped. On a complex task designed to trip up the model, consultants who used AI were nineteen percent less likely to reach the correct solution than those who did not. They had trusted a confident, wrong collaborator, and it cost them.

The experiment also surfaced distinct working styles. Some consultants behaved like “centaurs,” delegating clean subtasks to the model and handling the rest themselves. Others behaved like “cyborgs,” integrating the model into nearly every step. Both could succeed inside the frontier; the risk was concentrated in the moment the task crossed the boundary and the person did not notice.

The lesson is not that AI is unreliable in general. It is that the boundary is invisible from the inside. A model that is superb at one kind of reasoning can be fluent and wrong at an adjacent kind, and its fluency is precisely what makes the error persuasive. Getting smarter with these tools therefore requires an active relationship to that boundary: knowing what you are delegating, checking the parts that matter, and keeping your own judgment in the loop on the tasks where the model’s competence is uncertain. The capability gain is real, but it is conditional on the user’s calibration.

Cheap confidence and the erosion of checking

There is a subtler cost that shows up less in task scores than in habits of mind. In a 2025 study at CHI, Hao-Ping (Hank) Lee and colleagues surveyed 319 knowledge workers across 936 real examples of AI use at work. They found a consistent pattern: the more confident a worker was in the AI’s abilities, the less critical thinking they applied, and the lower their self-confidence, the more they deferred. Workers who trusted their own judgment checked the model more and reasoned harder about its output. Workers who trusted the model relaxed.

The study also describes a shift in what the work becomes. Instead of generating the analysis, the worker increasingly verifies, curates, and stewards output that the model produced. That can be a legitimate division of labor. It can also be a slow transfer of the skill of thinking itself to a system that never gets tired and never explains when it is guessing. If your job is to catch the machine’s mistakes, and you have stopped practicing the underlying skill, the checking degrades. The tool does not have to be wrong to make you worse. It only has to make the effort feel unnecessary.

Sorting demonstration from speculation

Because this subject invites hype in both directions, it is worth keeping three categories separate.

The first is demonstrated science. We have randomized, replicated evidence that a well-designed AI tutor can accelerate learning in a specific course, that an unguarded chatbot can harm unassisted performance, that AI assistance helps inside a competence frontier and hurts outside it, and that trust in AI correlates with reduced checking. These are measured effects with stated conditions. They are strong precisely because they are narrow.

The second is plausible engineering. It is reasonable to expect that these results can be extended — that better tutoring designs, adaptive difficulty, and reserved practice tasks will keep producing gains across more domains. This is an engineering bet. It is well motivated, it is being tested, and it is not yet established as general. When a company claims that its assistant will raise your capability, it is asserting something in this category, not the first.

The third is speculation. Claims that current systems meaningfully raise general intelligence, that they rewire the brain, or that they constitute a step toward direct neural enhancement belong here. They may turn out to be right in some form. They are not what the studies above show, and presenting them as settled is the most common failure in popular coverage of this topic. For the engineering of neural interfaces proper, the honest position is that the field is far from consumer implants, and that the near-term story about human capability is the far more mundane one told by the physics course and the Turkish classroom.

Making the gains stick

If the evidence says the benefit depends on design, then design becomes the practical question for anyone who wants to use AI to actually become more capable rather than merely more productive this afternoon. A few principles follow directly from the research.

Keep the machine out of the part you want to strengthen. If you want to reason well about a domain, do your own retrieval before asking, and ask for hints rather than answers. The Bastani result is close to a controlled demonstration that answer-giving inverts the benefit. Using a model as a Socratic partner, one that asks you why and pushes back, keeps the cognitive work where it belongs.

Reserve tasks the tool never sees. The clearest way to know whether AI is helping you learn is to have a set of evaluation problems you solve unassisted and score honestly, before and after. This is the transfer condition that separates real learning from performance on the task at hand. If your unassisted performance is not rising, you are being assisted, not improved.

Alternate assisted and unassisted practice. The Harvard and Turkish studies point the same way: benefit comes from a cycle where the tool scaffolds a stretch of practice and then steps back. Pure assistance builds dependence; pure struggle without feedback is slow. The productive pattern uses the model to accelerate the hard part — finding the edge of your competence, generating varied practice, supplying timely correction — and then tests you without it.

Watch your confidence. The CHI survey found that over-trust predicts less checking. A simple countermeasure is to treat fluency as a warning sign rather than a guarantee, and to make verification cheap and habitual on exactly the tasks where being confidently wrong is expensive.

Keep a record you can read against. Because the model can make you feel more capable whether or not you are, the only reliable evidence is a log of your own unassisted performance over time, kept separate from the sessions where you use the tool. This is the same discipline the Turkish researchers imposed by testing without the chatbot, and it is the one most people skip because it is unpleasant. A short set of problems solved cold, once a month, tells you more than any amount of satisfaction with the tool.

Creativity: more ideas, a narrower range

One more study complicates the picture in a useful way. Doshi and Hauser, publishing in Science Advances in 2024, gave writers either AI-generated story ideas or none, and then assessed the resulting stories. Access to ideas made individuals more creative — their stories were judged more novel and better written — but it also made the collective output more similar. The AI-assisted writers converged on the same themes. Individual capability rose while the diversity of the group fell.

This is a real trade and not a slogan. If the goal is to help a single person generate more and better options, AI does that. If the goal is a portfolio of genuinely different ideas, sameness is a cost. A person who wants to use these tools to think better has to decide which they are optimizing, and sometimes the answer is to use the model first and then deliberately diverge from it.

What would change the verdict

The question in the title resists a permanent answer because the tools and the studies are both moving. What would count as genuinely decisive evidence that AI is making people smarter rather than merely better equipped?

First, transfer. Gains should appear on unassisted tasks and on tasks the system never saw, measured after meaningful delays. Performance during a session is the weakest signal available.

Second, independence. Skill should not degrade when the tool is removed. The Turkish exam result is the model for what to check, and it should be checked for every claim of enhancement.

Third, breadth. Effects should hold beyond a single course, a single subject, and a single model version. The current strongest findings are narrow by design.

Fourth, calibration. Users should become better at knowing when to trust the model and when to override it — the frontier sense that the consultants lacked. A capability that raises output while destroying judgment is not capability in any durable sense.

Current AI can already clear the first bar in specific settings and the second bar when the interaction is designed for it. That is genuinely good news, and it is enough to justify using these tools deliberately today. What no study yet shows is a general, lasting increase in human intelligence delivered by software alone. The plausible version of the near future is less dramatic and more useful: not smarter people by default, but people who can be measurably more capable when they treat the model as a sparring partner rather than an oracle. The difference between those two futures is a set of design choices, and, for now, they are choices the user makes every time they decide whether to ask for the answer or for the help.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home