For most of the history of computing, the channel between a human intention and a machine action has been muscle. Fingers on keys, a hand on a mouse, a voice into a microphone. The model on the far side of that channel has changed beyond recognition in three years. The channel itself has barely changed in forty. That mismatch is why “from ChatGPT to BrainGPT” has become shorthand for a genuine research program, and why the phrase is so easy to overload with meaning it has not earned.
The story has three layers, and keeping them apart is the whole task. Brain-to-text decoding that has been demonstrated in human beings, mostly in patients who have lost the ability to speak or move. Plausible engineering, meaning non-invasive sensors, wearable devices, and interfaces that might someday serve a healthy user. And speculation, the seamless thought-to-AI channel that shows up in product announcements and press write-ups but does not exist as a working system. A single sentence can cross all three layers in one clause. Slowing that crossing down is the point of what follows.
The interface nobody notices
Compare the speeds at which a person can move information out of their head and into a machine. Ordinary typing runs at roughly forty words per minute for an average adult, and lower when the task involves unfamiliar vocabulary or careful thought. Conversational speech runs at about 160 words per minute, which is why voice input feels so much faster than a keyboard when it works. Both numbers hide a bottleneck upstream of the hands and the mouth: the rate at which a person can decide what to say. Most of the time spent composing a sentence is not spent moving fingers.
A brain-computer interface proposes to remove the motor stage entirely and read the intention directly. The premise is that the cortex already represents a planned movement, or an intended word, before the muscles act, and that a decoder can recover it. That premise is not fantasy. It has been tested, in people, with results specific enough to reason about.
What the clinic has actually built
The clearest evidence comes from a single participant reported in the New England Journal of Medicine in 2024 (Card et al., 2024). The patient was a 45-year-old man with amyotrophic lateral sclerosis, tetraparetic and severely dysarthric, who could no longer make himself understood by anyone but his regular care partner. His speech was intelligible only to expert listeners, at a rate of 6.8 correct words per minute; his typing with a gyroscopic head mouse ran at 6.3. Surgeons implanted four microelectrode arrays, 256 electrodes in total, into his left ventral precentral gyrus, a region that coordinates the movements of speech.
On the first day of use, twenty-five days after surgery, the system decoded his attempted speech with 99.6 percent accuracy using a 50-word vocabulary, after thirty minutes of calibration. On the second day, with 1.4 additional hours of training, accuracy was 90.2 percent over a 125,000-word vocabulary. With more data still, the team reported 97.5 percent accuracy sustained across 8.4 months, with the participant holding self-paced conversations at roughly 32 words per minute for more than 248 cumulative hours.
Those numbers deserve to be read slowly. The clinically meaningful one may be less the peak accuracy than the phrase the authors chose: useful function beginning on the first day of use. Before this work, speech BCIs typically required many hours of recording before they became usable at all. Rapid calibration is what turns a research demonstration into something a patient might actually live with.
It is not the only result worth keeping. A 2023 Nature paper from a different group reported a speech neuroprosthesis that reached a 9.1 percent word error rate on a 50-word set and 23.8 percent on a 125,000-word set, at about 62 words per minute, roughly three and a half times the previous record (Willett et al., 2023). For scale, conversational English runs near 160 words per minute, so even a fast speech BCI is still moving at about a third of natural pace. A separate line of work decodes handwriting rather than speech: two arrays in the hand area of premotor cortex let a paralyzed participant produce about 90 characters per minute at 94 percent raw accuracy, exceeding 99 percent with autocorrect (Willett et al., 2021). That is roughly 18 words per minute against about 115 characters per minute for a smartphone user of the same age. The device closes something like half the gap to an able-bodied thumb.
The most recent work pushes past text into voice. A 2025 study in Nature Neuroscience used high-density surface recordings to drive a continuously streaming speech synthesizer that output the participant’s pre-injury voice, decoding in 80-millisecond increments so that the rhythm of conversation survived (Littlejohn et al., 2025). Speech, in these systems, is becoming not just text but sound.
Every one of these results shares the same boundary conditions, and they matter as much as the headline numbers. Each came from a small number of participants, often one. Each used surgically implanted electrodes, in most cases penetrating the cortex. Each ran inside a clinical research protocol with trained staff in the room. These constraints are the conditions that made the results possible at all.
The non-invasive divide
If a person cannot or will not have electrodes placed in the brain, the picture changes sharply. Non-invasive recordings, electroencephalography (EEG) and magnetoencephalography (MEG), measure the summed activity of large populations of neurons from outside the skull. The signals are far weaker and blurrier, and the decoders that work on implanted arrays generally do not transfer.
The most optimistic non-invasive result to date decodes speech perception rather than production, and does so in healthy volunteers rather than patients. A 2023 paper in Nature Machine Intelligence trained a contrastive model across four public datasets and 175 volunteers recorded with MEG or EEG while they listened to speech (Défossez et al., 2023). From three seconds of MEG signal, the model identified the correct audio segment out of more than a thousand possibilities with up to 41 percent top-1 accuracy on average, rising above 80 percent for the best participant. The same approach applied to EEG reached only 17.7 to 25.7 percent top-10 accuracy. The gap between the sensors is the story: MEG is a different class of measurement, and it requires a magnetically shielded room and a motionless participant. It is a laboratory instrument, not something a person wears.
The closest thing to a non-invasive “BrainGPT” is a project called Brain2Qwerty, developed by Meta with academic collaborators, that decodes typed sentences from EEG or MEG in healthy volunteers (Lévy et al., 2025). With MEG, the system reached a character error rate of about 32 percent on average and 19 percent for the best participant, and could perfectly decode some sentences it had never seen. With EEG, the error rate rose to 67 percent. The project reported a scaling relationship in which more data improved accuracy in a predictable way, which is the strongest argument that the approach has room to grow rather than a hard ceiling.
Set that comparison directly beside the invasive numbers. An implanted array in a paralyzed patient reached 97.5 percent accuracy at a conversational 32 words per minute and held it for months. A non-invasive MEG system reading a healthy volunteer’s typing reached a character error rate between roughly a fifth and two-thirds, in a shielded room, in a laboratory. The distance between those two columns is the most important fact in the field, and it is routinely flattened by coverage that treats “brain decoding” as one capability.
Why the name “BrainGPT” oversells the moment
The label BrainGPT borrows the cultural weight of a product that answers arbitrary prompts and applies it to devices that decode a narrow, rehearsed task. The borrowed authority does real work on the reader, so it is worth naming precisely what is being borrowed.
Large language models are general: the same weights handle a legal question, a poem, a repair manual. Today’s neural decoders are not. They are trained for a specific person, a specific task, and often a specific vocabulary. Combining a decoder with a language model can improve output, because a language model can clean up ambiguous phonemes using context, in the same way autocorrect fixes a stray keystroke. But that is a language-model assist, not evidence that the brain signal is being read generally. When a system seems to know what a person meant because the model predicted the likely next word, the improvement comes from statistics, not from a better sensor.
The gap between announcement and demonstration is visible in the commercial layer. In January 2026, OpenAI announced that it was participating in the seed round of a startup called Merge Labs (OpenAI, 2026). The stated ambition is to build brain-computer interfaces that connect at “much higher bandwidth” by combining biology, devices, and AI, with co-founders drawn from academic work on ultrasound and neural engineering. What OpenAI described was a long-term mission and a research lab, not a product. The language of the announcement, speaking of molecules instead of electrodes and deep-reaching modalities like ultrasound that avoid implants into brain tissue, describes an intended direction. No working high-bandwidth, non-invasive, at-home neural interface has been demonstrated in humans. The correct way to read the post is as a statement of intent by a well-funded team, which is a meaningful signal about where capital is flowing and no evidence at all about what is technically possible today.
This is where the three-layer distinction earns its keep. The demonstrations live in the clinic. The engineering lives in the laboratory and the startup pitch. The speculation lives in the sentence claiming a person will soon think at their AI and the AI will think back.
The bandwidth underneath the question
Zoom out from any single paper and a structural fact appears. The brain moves far more information internally than any interface to it can carry. A sensory channel like vision delivers data at a rate that dwarfs language, while spoken language transmits on the order of a few dozen bits per second. The exact figures belong in the next article, but the shape of the asymmetry is the point here: speech is a narrow pipe attached to a wide one, and it is the narrow pipe that has carried human cooperation for as long as there have been humans.
That asymmetry is why the framing underneath BrainGPT tends to skip a prior question. If two people can reach the same model, are they equally augmented? A faster channel between mind and machine is a different claim from a smarter machine, and the two get bundled together whenever someone predicts that the interface is about to disappear. What a decoder at 32 words per minute changes is how quickly a person can convert an already-formed thought into text. It does not change how quickly the thought forms, how much of it the person can hold at once, or whether the person has anything worth saying. Those constraints sit on the human side of the port, and no electrode removes them.
What would have to be true
Suppose the goal is genuine, general, consumer neural communication: not a clinical prosthesis, but a channel a healthy person uses the way they use a keyboard. A short list of conditions would have to hold, and checking them against the current state is a useful discipline.
A non-invasive sensor would have to work outside a shielded room, while the person moved and spoke. MEG fails on both counts today. EEG survives motion and ordinary environments, but its error rate in the typing task was 67 percent, and it does not approach the implanted numbers even in perception decoding. A wearable optical or ultrasound sensor that reaches deep cortex at high resolution would be a genuine breakthrough; several research programs are chasing exactly that, and none has yet produced a usable device.
The decoder would have to hold up across sessions without hours of recalibration. Here the invasive work has made real progress: the 2024 patient’s system remained stable across months with periodic maintenance. The non-invasive work is more fragile, with signals that drift across sessions and individuals, which is why models often need a participant-specific layer.
The system would have to be safe under daily use for years. Penetrating electrodes raise long-term questions about tissue response that a lifetime of consumer use would magnify. Non-invasive methods avoid that risk, which is precisely why the field keeps returning to them despite their lower resolution.
And the person would have to want it for reasons that survive scrutiny, not because an employer requires it or because the alternative is unaffordable. That condition is not an engineering problem, and it is the one most likely to be skipped.
Where the risk sits
Mind reading is the wrong worry; the signals decoded today are task-bound and require active cooperation. The clearer danger is inference at scale. A device that decodes intended speech also produces a record of everything the person intended to say, and neural data is unlike other data because it is generated involuntarily and is hard to change. A password can be reset; the shape of a person’s cortical activity cannot. The governance questions arrive before the technical ones are settled, and the ethics literature is already behind. A 2025 scoping review of closed-loop neurotechnology found that only one of sixty-six clinical studies included a dedicated ethical assessment (Haag et al., 2025).
The second danger is the objective, and it is the same danger that runs through the rest of this series. A decoder that serves a patient’s own goal of being understood is an assistive technology. The same hardware pointed at an objective the user did not choose, whether engagement, productivity, or compliance, becomes something else. The channel does not care which direction the intention flows. The design and the governance do.
Where this leaves us
The road from ChatGPT to BrainGPT is real, and it runs through the clinic first. In patients who have lost speech, direct brain-to-text and brain-to-voice interfaces have moved from experiment toward something close to a daily tool, with accuracy, speed, and calibration that would have seemed implausible a decade ago. That is the demonstrated layer, and it is a substantial human achievement.
The second layer, a non-invasive interface good enough for a healthy person to use casually, is genuinely being worked on and is not close. The third layer, the seamless thought-to-AI channel, is a projection. The most useful posture toward the whole field may be to ask, of any claim, which layer it belongs to, and to notice how often the answer is not the one the sentence implied.
Sources and further reading
- Card et al., “An Accurate and Rapidly Calibrating Speech Neuroprosthesis,” New England Journal of Medicine, 2024
- Willett et al., “A high-performance speech neuroprosthesis,” Nature, 2023
- Willett et al., “High-performance brain-to-text communication via handwriting,” Nature, 2021
- Littlejohn et al., “A streaming brain-to-voice neuroprosthesis to restore naturalistic communication,” Nature Neuroscience, 2025
- Défossez et al., “Decoding speech perception from non-invasive brain recordings,” Nature Machine Intelligence, 2023
- Lévy et al., “Brain-to-Text Decoding: A Non-invasive Approach via Typing,” arXiv, 2025
- OpenAI, “Investing in Merge Labs,” 2026
- Haag et al., “Ethical gaps in closed-loop neurotechnology: a scoping review,” npj Digital Medicine, 2025
Loading comments…