Put two people in front of the same frontier model. Give them identical access, identical billing, identical latency. Are they now equally augmented? The intuitive answer is yes, because the capability is the same. The more careful answer is no, because the model is only one component of a system that includes the person, and the person brings constraints the model cannot remove. The most underrated of those constraints is bandwidth: the rate at which a person can get an intention, a piece of context, or a correction into the system and get a useful result back out.
This is a claim that has to be handled carefully, because it sits between a measurable fact and a loose metaphor. The measurable fact is that human communication channels carry remarkably few bits per second. The metaphor is that a person is a “low-bandwidth device” who just needs a faster port. The first is real and well-studied. The second smuggles in an assumption that the bottleneck is entirely on the transport side, when a large part of it sits in the machinery that decides what to transmit.
The numbers that make the argument concrete
Start with the sensory side, because that is where the numbers are largest and cleanest. In a 2006 study in Current Biology, a team recorded from guinea pig retina on a multi-electrode array and measured how much information the optic nerve actually carries (Koch et al., 2006). They estimated that about a hundred thousand ganglion cells transmit on the order of 875,000 bits per second. Scaling to the roughly one million ganglion cells of the human retina, they concluded that the eye sends data to the brain at roughly 10 million bits per second, about the rate of an Ethernet connection.
That is the input channel. Now the output channel, language. A 2019 paper in Science Advances analyzed recordings from 170 native speakers across 17 languages and found something striking (Coupé et al., 2019). Languages vary enormously in how much information each syllable carries, and in how fast their speakers talk, but the two properties compensate: fast-talking languages pack less information per syllable, dense languages are spoken more slowly. Across the whole sample, the rates converge on roughly 39 bits per second. That number is the rate at which one person transmits meaning to another through speech.
Set the two figures side by side. The eye pulls in on the order of ten million bits per second. Speech pushes out about forty. Visual input is roughly two hundred thousand times faster than verbal output. Even allowing for the fact that these are different kinds of measurement, comparing a raw sensory channel against a linguistically coded one, the asymmetry is enormous, and it is the asymmetry every interface to AI has to live inside.
What the asymmetry implies
The tempting reading is that humans are built for input and starved on output, so the obvious move is to widen the output channel. That reading is partly right and partly a trap, and separating the two is where the interesting analysis begins.
The output channel is narrow partly because language is a compression scheme. A speaker does not transmit a picture of a scene; they transmit a few dozen bits per second that, combined with what the listener already knows, reconstructs the scene. The compression works because the listener shares context: a lifetime of experience, a common culture, a model of the speaker, an ongoing situation. Most of the information in a conversation is not in the words. It is in the prior knowledge the words point at.
This is why the intuition that a faster port will make thinking faster runs into trouble. If the bottleneck were purely the transport rate, then giving a person a direct neural channel at ten million bits per second would translate into a ten-million-bit-per-second conversation. But conversation is not a bulk file transfer. The reason speech is slow is not that the mouth cannot move faster. It is that both parties can only hold and process so much novel meaning at once. Widening the pipe below a bottleneck that sits upstream of the pipe does not increase throughput.
The bottleneck above the port
The upstream constraint has a name in cognitive psychology: working memory. The classic figure is Miller’s seven plus or minus two, from 1956, but later work tightened it. In a 2001 review in Behavioral and Brain Sciences, Nelson Cowan argued that when chunking and rehearsal are blocked and capacity is measured cleanly, the limit is closer to four chunks, not seven (Cowan, 2001). A person can hold roughly four independent items in the focus of attention at a time.
Four items is not a transport limit. It is a limit on what can be held active simultaneously, and it applies whether the items arrive through a keyboard, a voice, or an electrode. When a person uses a language model to think through a problem, the constraint that shapes the session is often not how fast they can type but how many distinct threads they can keep in mind while typing. A faster interface shortens each round trip. It does not increase the number of threads a person can track.
This is the crux of why “bandwidth may matter more than intelligence” is a claim worth stating precisely rather than repeating loosely. Two people with the same model are not equally augmented mainly because of differences in what they can do with the model’s output, and those differences include working memory, domain knowledge, prior skill at decomposing problems, and the metacognitive habit of noticing when a line of reasoning has gone wrong. A faster channel helps a skilled operator more than it helps a novice, because the skilled operator can absorb the result and has somewhere to put it.
Why the metaphor still earns its place
None of that dissolves the bandwidth argument; it relocates it. There is a real sense in which the model is constrained by the human operator’s throughput, and there are real ways that improving throughput changes outcomes.
Consider latency and iteration count. A person who can complete twenty ask-refine cycles in the time another completes five will, other things equal, converge on a better result sooner. Voice input is faster than typing and frees the hands; the gap between roughly forty words per minute typed and 160 spoken is a fourfold increase in raw output rate. That is why voice changes how people use these systems, even without any neural hardware. The gain is not that each sentence is better but that more sentences fit in a fixed span of attention.
Consider what the interface lets a person include. A prompt typed on a phone strips away detail because the input cost is high. A prompt spoken aloud, or drafted with a tool that has access to the user’s documents and screen, can carry more context at the same effort. Here bandwidth is not only speed; it is richness per unit of user effort. A person whose interface lets them gesture at a spreadsheet, past a screenshot, and reference a file by name has effectively widened the channel without changing the biological port.
And consider the feedback direction. Bandwidth runs both ways. A system that shows its work, surfaces uncertainty, and lets a person steer mid-stream gives that person a channel into the model’s process. A system that returns a single opaque answer and stops narrows the person’s ability to correct it, no matter how fast they can type. The design of the return path can matter more than the design of the input path, and it is usually cheaper to improve.
The daily-life evidence
The clearest human evidence that the interface shapes the outcome predates AI by decades. It comes from research on transactive memory, the practice of storing information partly outside the mind and partly in the people and tools around us. In a 2011 study in Science, participants who expected to be able to look information up later remembered the information itself less well but remembered where to find it more reliably (Sparrow et al., 2011). In one experiment, participants recalled folder locations better than they recalled the statements stored in those folders, with mean recall of roughly 0.49 for locations against 0.23 for the statements themselves.
That finding is usually cited as a warning about dependence on search engines. It is also a description of how a person offloads part of a task to a channel. The mind treats an accessible external store as an extension of memory, and it reallocates its own limited capacity toward knowing where the knowledge lives. Read that as a design brief: the more reliably a channel can be queried, the more a person will delegate to it, and the more the person’s competence comes to depend on the channel remaining available and trustworthy. A wide, dependable channel makes a person faster and, in a specific sense, thinner. Both effects are real, and they trade off against each other.
Where intelligence and bandwidth actually separate
It is worth being concrete about the difference, because the two are bundled together in most forecasts. Intelligence is a property of the model: how well it generalizes, how reliably it reasons, how much context it can hold. Bandwidth is a property of the loop between the model and the person: how much of the person’s intention reaches the model, how much of the model’s result the person can absorb, and how quickly corrections flow in both directions.
A jump in model intelligence raises the ceiling for everyone. A jump in loop bandwidth raises the ceiling only for people who have something to put into the loop and the capacity to receive what comes out. That is why the same model, deployed to a skilled researcher and to a first-time user, produces outcomes that diverge more than the raw capability difference would predict. The model is constant; the loop is not.
This also explains a persistent puzzle in how AI adoption spreads. Capability improves on a predictable curve, but the realized productivity gain varies wildly across individuals and organizations. The variation tracks how well the human side of the loop is engineered: whether the interface fits the task, whether context is easy to supply, whether corrections are cheap, whether the person has the judgment to know when to trust the output. None of those are model problems.
What a wider channel would and would not buy
Suppose the interface research succeeds and a healthy person can send and receive at orders of magnitude more than forty bits per second. What changes? The honest answer is: the parts of the task that were limited by serialization and by manual transcription. A person who thinks faster than they can type benefits directly. A person working through a large document benefits from being able to indicate a location in a structure rather than describing it in words. A person juggling many context sources benefits from an interface that gathers them automatically.
What does not change is the working-memory ceiling, the need to know what is worth asking, and the requirement to verify what comes back. A direct channel does not enlarge the focus of attention; if anything, a faster flood of output could overwhelm a system tuned for forty bits per second of language. A very wide channel fails less by being useless than by exceeding the human’s ability to filter, producing something closer to noise. The design problem after the bandwidth jump is the same design problem before it: deciding what deserves the person’s limited attention.
There is also a subtlety about what “bandwidth” means for intention. Much of what a person wants to convey is not explicitly represented yet at the moment of transmission. A thought is often not a finished sentence waiting to be sent; it is a direction, a discomfort with a current framing, a sense that something is missing. Language has to externalize that into structure before it can travel. A faster port does not eliminate that translation step. It moves data faster between two points that are still doing the work of turning vague promptings into transmissible content.
The failure modes that come with the channel
The risk in widening the human-machine channel is not the channel itself. It is where the widening is pointed, and who controls the endpoints.
The first failure mode is dependence that becomes invisible. If a person offloads memory, judgment, and reasoning to a channel that is always available, the skills that used to run without the channel can atrophy, and the person may not notice because the channel keeps producing acceptable results. The transactive-memory research is the template: the mind adapts to the tool, and the adaptation is efficient until the tool is gone. A responsible system makes the person more capable with it than without it, rather than more capable only while plugged in.
The second failure mode is asymmetric channel design. A company can measure engagement from its side of the loop more finely than a user can measure their own benefit. If the return path is engineered to capture attention while the input path is engineered to be frictionless, the user is optimized, not augmented. The bandwidth that matters, in that case, flows toward the institution.
The third is the objective problem, and it is the same one that runs through every article in this series. A wide channel serves whatever goal is loaded into it. A channel that helps a person pursue their own goals widens their range of voluntary action. A channel that pursues a goal the person did not choose narrows that range while feeling convenient. Bandwidth does not decide which of these it is; the governance around the endpoints does.
What to watch
For any interface claim, ask where the constraint actually sits, rather than how wide the pipe is. If the current bottleneck is typing speed, then voice, better input methods, and context-aware tools address it. If the bottleneck is working memory, no interface removes it, and the relevant engineering is chunking, summarization, and external scaffolding that helps a person hold a line of reasoning without losing it. If the bottleneck is judgment, the answer is training and practice, not hardware.
It also matters whether the interface research and the model research move at the same pace. The models are improving faster than the human side of the loop, which means the marginal gain from a smarter model is shrinking relative to the marginal gain from a better-fitted interface. That is the practical version of the article’s thesis: at the frontier, the binding constraint is increasingly the human channel, not the machine behind it.
Two people with the same model will diverge, and that divergence is the best evidence for the claim. Measuring it is a way to keep the claim honest rather than rhetorical. The person who can hold four threads, supply context cheaply, and correct the system in flight will outperform the person who cannot, using a model neither of them owns or understands. Bandwidth, like attention, is a resource that distributes unequally even when the intelligence on offer is identical.
Sources and further reading
- Koch et al., “How Much the Eye Tells the Brain,” Current Biology, 2006
- Coupé et al., “Different languages, similar encoding efficiency: Comparable information rates across the human communicative niche,” Science Advances, 2019
- Cowan, “The magical number 4 in short-term memory: A reconsideration of mental storage capacity,” Behavioral and Brain Sciences, 2001
- Sparrow, Liu & Wegner, “Google Effects on Memory: Cognitive Consequences of Having Information at Our Fingertips,” Science, 2011
- Défossez et al., “Decoding speech perception from non-invasive brain recordings,” Nature Machine Intelligence, 2023
- Card et al., “An Accurate and Rapidly Calibrating Speech Neuroprosthesis,” New England Journal of Medicine, 2024
Loading comments…