In 2023, a research team reported in Nature that a volunteer with advanced ALS — a person who could no longer speak intelligibly — used an implanted brain-computer interface to generate text at roughly 62 words per minute, about three and a half times the previous record for a speech neuroprosthesis. With a fifty-word vocabulary the system’s word error rate was close to 9 percent. With a 125,000-word vocabulary it climbed to about 24 percent (Willett et al., 2023).
Hold those two numbers side by side, because the discipline this whole subject demands lives in the space between them. The first looks like a miracle of restored communication. The second shows how fast accuracy collapses once the decoding problem gets harder. Both are true at once, and neither one is mind reading.
The question in the title is not rhetorical. Today’s AI can extract genuine information from neural activity — more than most people would guess, and far less than most headlines imply. Sizing that gap correctly requires looking at how decoding actually works, at what each prominent study measured, and at the constraints that made its result possible. One pattern repeats across every method and every lab: the more tightly a task constrains what a person might be thinking, and the more directly a sensor touches the brain, the better the decoder performs. Loosen the constraints and the performance falls apart.
What a decoder is doing
A neural decoder does not read thoughts. It estimates a label from a signal. The signal is some physical measurement of brain activity — voltages from electrodes, magnetic fields from magnetoencephalography, blood-oxygen changes in functional MRI. The label is whatever the experimenter has decided to treat as the target: a word, a letter, whether a person saw a house or a face, a scaled mood score.
The relationship between signal and label is learned, not given. A model is trained on examples where both are known — recordings taken while the person performed a known task — and then asked to produce the label from the signal alone. Everything about how well it works depends on choices made before the model ever runs: which signal was recorded, how the task was framed, how many candidate answers the model may choose among, and how much the person cooperated.
Those choices are the reason a decoder can look nearly perfect in one study and useless in another without anything mystical happening. Decoding a person’s intended word from a closed set of fifty is a different problem from decoding it from a 125,000-word vocabulary. Reporting a “top-1 accuracy” means the model picked the single best candidate; reporting “top-10” means the correct answer was somewhere in ten guesses, which is a far easier bar. A number without its task is not a fact about technology. It is a fact about an experiment.
The invasive ceiling
The strongest results come from electrodes placed inside the skull, in direct contact with cortex. The Willett study is the clearest benchmark. Four microelectrode arrays were implanted in the motor cortex of one participant with ALS, and the group’s model converted attempted speech into text continuously. The headline figure — 62 words per minute — is fast compared with earlier systems, though still slower than ordinary conversation, which runs closer to 160 words per minute. The error rates tell the real story of the trade-off: about 9 percent when the vocabulary is small, about 24 percent when it is large.
Those numbers describe an extraordinary result under extraordinary conditions. The participant had undergone brain surgery. The recordings came from hundreds of neurons that a surgeon had reached by penetrating the cortex. The task was speech itself, the person was actively trying to talk, and the model was tuned to that one individual’s neural patterns over many sessions. This is not a criticism of the science; it is a description of what the science requires. Invasive approaches reach the signal because they eliminate the skull’s attenuation, and they achieve their accuracy because the problem is bounded and the participant is a willing, cooperative subject.
A decoder that restores a voice to someone who has lost one is an unambiguous good. The mistake is to read the same numbers as evidence about reading anyone’s thoughts, anywhere, without surgery or cooperation. Nothing in the invasive result supports that leap.
Decoding speech without opening the skull
The non-invasive versions are where the popular imagination runs ahead of the evidence. The most cited recent effort is Brain2Qwerty, a project from Meta’s research group that decodes typed sentences from brain recordings taken outside the skull. Its first version recorded magnetoencephalography (MEG) and electroencephalography (EEG) from thirty-five healthy volunteers while they typed briefly memorized sentences (Lévy et al., 2025; project page). On average, MEG produced a character error rate of about 32 percent and EEG a far worse 67 percent, with the best participants reaching roughly 19 percent.
Read the fine print. The task was typing, not free thought, and the authors’ own error analysis found that decoding leaned on the motor signals generated by the act of typing as well as on higher-level language. Even the best participant mis-typed roughly one character in five. The system works because the problem is narrow and the cues are rich. A later version of the project reported a substantial gain — a mean near 61 percent word accuracy and up to about 78 percent for the best participant — but its own summary notes that the error rate remains too high for everyday use, and the sensor is still a room-sized MEG scanner. You cannot wear it, and you cannot move much while it records.
A different 2023 study took a still easier route: perception rather than production. Rather than decoding what a person was saying, it decoded what they were hearing. Across 175 volunteers and four datasets, MEG models identified the correct speech segment from more than a thousand alternatives with top-1 accuracy of about 41 percent on average and above 80 percent for the best participants, while EEG lagged far behind at roughly 18 percent top-1 and 26 percent top-10 (Défossez et al., 2023). The study included speech segments absent from training, a useful test of generalization within the decoding task. That result should not be confused with effortless generalization to a new person or a new kind of thought. It is also, by design, a passive task. The brain was doing something the experimenter had arranged for it to do, and the decoder matched the output to a menu of options it had already been given.
That contrast — perception versus production, recognition versus generation — is one of the deepest divides in the field. Recognizing which of a thousand known sounds a person heard is a fundamentally easier computational problem than generating, from scratch, the sentence a person intends to say. Reporting a perception result as if it were a production result is the most common way coverage overshoots.
Gist, not transcript
Functional MRI offers the most vivid demonstrations and some of the most easily misread ones. In 2023, a group showed that it could reconstruct the meaning of continuous language from non-invasive fMRI recordings (Tang et al., 2023). Three participants listened to stories, imagined stories, and watched silent films while their blood oxygenation was recorded; after about sixteen hours of training data per person, a model recovered the gist of what they were hearing or imagining, well above chance.
What the decoder produced was paraphrase, not transcript — the broad sense of a passage, not its words. It required the participant’s cooperation both to train the model and to apply it, since the model must be calibrated to the individual brain. And the underlying signal is slow: the hemodynamic response it measures lags neural activity by several seconds, which limits temporal precision and delays the reconstructed output. The achievement is that meaning left a readable trace in a non-invasive signal. The limit is that meaning is all it recovered, and only with the person’s active help.
Image reconstruction follows a similar shape. In 2023, researchers combined fMRI with a latent diffusion model, the kind of generative AI behind text-to-image tools, to reconstruct images a person had viewed (Takagi & Nishimoto, 2023). The results are striking to look at. They are also blurry, approximate, and dependent on a model trained on thousands of images from one person at a time. The reconstruction resembles the original in layout and category — a red object where a red object was — but does not reproduce it. Contemporary coverage noted the same two facts that matter here: the requirement of a per-person model and the vagueness of the output. Neither is what one would expect from mind reading, whatever the images look like.
Reading a state, not a sentence
Not every decoder targets language or images. Some target something more like a mood or an internal state, and here the invasive work is again the strongest. In 2018, a team reported that moment-to-moment mood variation could be decoded from intracranial recordings in seven people who already had electrodes implanted to locate epileptic seizures (Sani et al., 2018). Participants intermittently rated their mood on a tablet, and a model learned to predict those ratings from coordinated activity across distributed brain regions, mostly in limbic areas.
This is a remarkable finding, and it is worth being precise about why. Mood is subjective and diffuse; that it left any decodable trace at all, in a coordinated network rather than a single spot, is a real discovery. But the conditions were narrow. The recordings were invasive. The sample was seven people, recruited over years because the data are so hard to collect. The model was built for each individual. Nothing here shows that a consumer device can read a person’s feelings from the outside, and the distance between “an implanted array in a hospital can estimate a scaled mood score for one patient” and “a headband can tell how you feel” is the entire distance this article is about.
Why the headlines outrun the lab
The gap between a laboratory decoding result and unrestricted mind reading is not a matter of a few more years of engineering. It is structural, and it comes from at least five places.
Task specificity comes first. The strongest results belong to constrained tasks, although some generate language across a large vocabulary. Speech decoding knows the language and the vocabulary. Perception decoding knows the menu of stimuli. Image reconstruction knows that an image is coming. A person thinking freely, in whatever language, about whatever they like, offers a target set that is effectively infinite, and no present decoder has been asked to search it.
Cooperation comes second. Decoders are trained on a person doing a known thing, and most need that person to keep doing it. The fMRI language decoder had to be calibrated to each individual and applied while the person engaged with the task. A decoder that only works when its subject is actively helping is not an instrument of covert surveillance.
The sensor’s reach comes third. The most accurate methods are invasive. The non-invasive ones either need a stationary scanner, as with MEG and fMRI, or lose so much signal at the skull that accuracy drops sharply, as with EEG. EEG, the only genuinely portable electrical method, was the weakest performer in the studies above by a wide margin.
Generalization comes fourth. A model tuned to one person’s anatomy and neural patterns often fails on the next person, and models can drift as the brain changes with learning, fatigue, or time. Cross-day and cross-person stability remains an open problem, and it is the one that separates a demonstration from a product.
The fifth factor is the least appreciated: the difference between decoding a proxy and reading a thought. A decoder estimates the label the experimenter chose. When that label is a word from a known set, the mapping is tight. When the label is “mood” or “attention,” the decoder is estimating a proxy — a scaled self-report, a reaction time — and a model can track a proxy accurately while remaining wrong about the person. That gap between the measurement and the meaning survives even a perfect classifier.
What would actually count
Because “can AI read the brain?” is a yes-or-no question with a misleadingly simple shape, it helps to state what a genuine demonstration would have to show. It would need to work across days and people without lengthy recalibration. It would need to handle targets the system was not told in advance. It would need to function without the subject’s active cooperation, in ordinary environments, with a sensor a person could plausibly wear. And it would need to hold up under replication, with reported error rates and adverse events, rather than surviving on a vivid demo.
No present system clears those bars, and none is close. The best decoders available today do their work with the participant seated, cooperating, and performing a task the experimenter designed, and their accuracy is evaluated against known task targets. How performance changes when those supports are removed must be measured separately for each method. That is why the demonstrable capability and the imagined one are separated by more than a few years of engineering: the missing pieces are not merely better algorithms but a fundamentally more permissive measurement situation.
The honest summary is that present-day AI can decode constrained, cooperative, task-bound information from neural activity with real skill — and cannot decode unconstrained thought at all. The two statements are not contradictory. They describe a technology that is genuinely powerful within carefully built walls and much weaker outside them.
Why the distinction still matters
None of this is a reason for complacency. The capabilities that do exist are consequential precisely because they are reliable within their domains, and domains have a way of widening. A decoder that restores speech is also a decoder that, in the wrong hands and under the wrong incentives, extracts information a person did not choose to give. The reason to be precise about current limits is not to quiet the concern; it is to aim the concern at the real target. Fear of a fictional all-purpose mind reader distracts from the concrete questions that already have answers to give: who owns brain data, who can compel its collection, and what a person’s rights are over the activity inside their own skull.
Those questions belong to a longer argument about neural data and human agency, and the series returns to them elsewhere. For now the technical point stands on its own. The brain can be decoded, narrowly and with help. It cannot yet be read, in the sense the word usually invites. Confusing the two — in either direction, hype or dismissal — is the mistake to avoid.
Sources and further reading
- Willett et al., “A high-performance speech neuroprosthesis,” Nature, 2023
- Lévy et al., “Brain-to-Text Decoding: A Non-invasive Approach via Typing,” arXiv, 2025
- Défossez et al., “Decoding speech perception from non-invasive brain recordings,” Nature Machine Intelligence, 2023
- Tang et al., “Semantic reconstruction of continuous language from non-invasive brain recordings,” Nature Neuroscience, 2023
- Takagi & Nishimoto, “High-Resolution Image Reconstruction With Latent Diffusion Models From Human Brain Activity,” CVPR, 2023
- Sani et al., “Mood variations decoded from multi-site intracranial human brain activity,” Nature Biotechnology, 2018
Loading comments…