It is easy to argue that artificial intelligence will change everything, and just as easy to argue that it will ruin everything. Both claims share a weakness. They are hard to falsify, and neither tells anyone what to do next. The more demanding question is whether there is a positive case that survives contact with evidence — one that names mechanisms, points to results that already exist, and is candid about what has not yet been earned.
The strongest version of that case is not about effortlessness. It is about scarcity. Several of the things that make a human life go well are rationed, not because they are rare in principle, but because they are expensive to deliver: the ability to understand a genuinely hard problem, to recover a function the body has lost, to learn from a patient teacher, to see the consequences of a decision before making it. Each of those has already been moved, modestly but measurably, by narrow systems that do one thing well. The interesting question is how far the movement can go, and what has to remain human for the gains to count as human gains rather than as a smoother form of dependence.
The shape of a serious positive case
A positive case is not a mood. It is a claim with a structure. It should say what capability is being added, to whom, through what mechanism, and at what cost. It should separate what has been demonstrated in a controlled setting from what is plausible engineering from what remains speculation. And it should be honest that the same capability which widens one person’s options can be used to narrow another’s.
Being strict about this structure is what separates optimism from marketing. The enthusiast points at a benchmark and infers a civilization; the skeptic points at a failure and infers a collapse. Neither inference is licensed by the evidence. AlphaFold’s accuracy on a protein-folding benchmark does not, by itself, predict a golden age of medicine. A chatbot’s confabulation does not, by itself, predict a ruined generation. The work is to hold the specific result in view and reason outward carefully, one step at a time.
The other requirement is that the benefit land on the person, not only on the institution that deploys the system. A tool that makes a company more productive while making a worker more replaceable is not, from the worker’s side, good news. This series keeps returning to a single test: after the system acts, is the person more able to understand their situation and choose within it, or less able? A strong positive case has to pass that test, not merely promise a larger number somewhere in the aggregate.
Evidence that already exists
Optimism does not have to wait for the future. Four domains already contain results that are real, published, and independently checkable. Each system is narrow. That narrowness is exactly what makes them trustworthy as evidence, because the claim being made is bounded.
Reading a shape we could not compute
For roughly fifty years, predicting a protein’s three-dimensional structure from its amino acid sequence was a benchmark problem in biology. The search space is astronomical — a small protein has vastly more theoretical shapes than there are atoms in the observable universe — and the experimentally confirmed structures in public databases number in the hundreds of thousands, against hundreds of millions of known sequences. In 2020, DeepMind’s AlphaFold2 changed the terms of the problem. The system, described in Nature (Jumper et al., 2021), predicted structures from sequence with accuracy that, in the biennial CASP assessment, approached experimental resolution for many targets.
The recognition followed the result. The Royal Swedish Academy of Sciences awarded the 2024 Nobel Prize in Chemistry to Demis Hassabis and John Jumper for protein structure prediction and to David Baker for computational protein design (Royal Swedish Academy of Sciences, 2024). The Academy reported that the model’s predictions had been used by more than two million researchers in 190 countries and had produced structural predictions for essentially all of the roughly 200 million known proteins.
The number of users matters more than the medal. A predicted structure is a hypothesis that a wet-lab experiment can confirm or refute far faster than a blind search. The historical bottleneck in structural biology was never curiosity; it was the cost, in months and money, of getting an answer. When an answer gets cheaper, more questions get asked. That is the honest shape of the benefit: not a machine that displaces scientists, but a machine that removes waiting, and thereby changes which experiments are worth running.
A patient tutor for people who never had one
The education results should most embarrass the pessimists, because they come from randomized trials rather than from advertising. In 2025, a controlled study in Scientific Reports tested a purpose-built AI tutor inside a Harvard physics course (Kestin et al., 2025). Students who studied with the AI tutor achieved larger learning gains than students in the active-learning comparison condition, and did so in less time. The measured effect was not small.
Read the caveats alongside the finding. The setting was a well-resourced university. The tutor was engineered deliberately around pedagogical constraints rather than being a general chatbot. The comparison was against a good live class, not against no instruction at all. The study does not show that any model teaches anything well; it shows that a carefully designed system can produce measurable learning, which is a narrower and more useful claim.
The larger possibility is distributional. The scarce resource in education is not information — that has been effectively free for a generation — but a patient, responsive tutor who can diagnose a specific confusion and adjust. If even a fraction of that resource can be replicated at low marginal cost, the gain lands hardest on the people who could never buy it: students in understaffed schools, adults retraining mid-career, anyone whose learning stalled because no one had the time to explain the missing step.
Giving back a voice
Some of the clearest moral wins are the least ambiguous. A 2023 study in Nature described a speech neuroprosthesis that decoded attempted speech from cortical activity in a person with severe paralysis and rendered it as text at roughly 78 words per minute, alongside a synthesized voice and an animated avatar (Metzger et al., 2023). The reported word error rate was about one word in four — useful, imperfect, and improving with each generation of the research.
There is nothing speculative about the direction of that benefit. A person who cannot speak has regained a channel to the world. The engineering is invasive, restricted to participants with surgically implanted electrodes, and years away from routine clinical use. But the value is not in doubt, and it is the kind of value that no aggregate productivity statistic captures: the return of an ability that most people never think to count.
A faster look at the sky
Weather forecasting is where the positive case is easiest to audit, because physics-based models already exist and every forecast can be scored against what actually happened. Google DeepMind’s GraphCast, reported in Science, learned to produce ten-day global forecasts from decades of reanalysis data and outperformed the standard physics-based baseline on the large majority of the variables tested, while running in under a minute on a single accelerator (Lam et al., 2023). A follow-up system, GenCast, extended the approach to probabilistic forecasting, generating a full ensemble that beat the leading operational ensemble forecast on the overwhelming majority of targets (Price et al., 2024).
A forecast that is both faster and more accurate is not glamorous. It is also, measured in lives and money, one of the largest improvements a scientific instrument can deliver, because so much of agriculture, shipping, aviation, energy trading, and disaster response is a wager on what the atmosphere will do next. When the wager gets better, the losses fall. That the gain arrives as a quiet percentage of improved skill rather than as a headline does not make it smaller.
Why science is the strongest lever
It would be a mistake to treat these as four separate curiosities. They share a shape, and the shape is the argument. In each case, a learned system was pointed at a problem where the relationship between inputs and outputs was real, stable, and hard for unaided human minds to hold at once: the folding of a chain into a shape, the mapping from a student’s confusion to the right next explanation, the translation of neural firing into intended speech, the evolution of a planetary atmosphere. In each case the system compressed a space too large to search by hand into a usable guess, and in each case a human or an experiment could then check the guess cheaply.
Science is where this lever is strongest, for a structural reason. Scientific claims are already built to be tested, so a fast generator of hypotheses feeds directly into a slow, rigorous filter that the community already trusts. The model proposes; the experiment disposes. That division of labor is unusually healthy, because it keeps the machine in the role of a source of suggestions and leaves the authority to confirm or reject with the physical world. An AI that guesses well accelerates the rate at which guesses can be made and eliminated, which is most of what scientific progress is.
There is a temptation to inflate this into the claim that AI will “solve science.” The evidence does not support that yet. The results above are real, but they are also narrow and, in several cases, domain-specific in ways that do not transfer. What they collectively show is that the lever is real and already available, not that it has been fully pulled.
The honest ledger
The temptation at this point is to add the four stories together and announce a golden age. That would be the precise error the structure was meant to prevent. The track record supports a narrower claim: for well-posed problems with abundant data and a clear score, learned systems can now match or exceed traditional methods, and the resulting capability can be placed in many hands at low marginal cost. That is genuinely large. It is not the same as curing disease, closing educational gaps, or governing well.
It is worth setting the most influential economic estimate next to the technical optimism, precisely because the two do not obviously agree. Daron Acemoglu’s task-based analysis of AI’s macroeconomic effects concluded that, using existing evidence on which tasks are exposed and how much cost they save, the plausible gain in total factor productivity over a decade is on the order of a fraction of a percent, and likely less once the harder-to-learn tasks are accounted for (Acemoglu, 2024). One can dispute the estimate. The point of citing it is to model what a disciplined positive case looks like: it survives being compared against a skeptical number instead of being asserted past it.
The reconciliation is not that the technology is weak. It is that capability and measured economic effect are different variables, moving on different clocks. A tool can transform how a specific kind of scientific work is done and still take many years to show up as national productivity, because adoption is slow, institutions resist reorganization, and the tasks that resist automation are the ones where context and judgment dominate. Honest optimism can hold both facts at once without flinching from either.
Prediction is not judgment
There is a subtler reason the gains are not automatic, and it has nothing to do with malfunction. It has to do with deference.
A system that is right more often than a person creates real pressure to trust it. In the moment, that pressure is often rational: if the model is better calibrated than you are, overriding it can look like arrogance. But repeated deference is a kind of training, and it runs in an unhelpful direction. The competence to know when a confident system is wrong is itself built by making calls and living with the consequences. A person who has outsourced every prediction may discover, on the day the model fails, that they no longer have the muscles to notice.
This is why the best designs leave friction in the places where friction is useful. They show the evidence behind a claim. They expose uncertainty rather than collapsing it into a single confident label. They allow the output to be questioned. And they keep a named human or institution accountable for decisions that carry consequences. None of that makes the instrument worse. It keeps the human in the role of someone who understands, rather than someone who merely complies.
The distinction this series keeps drawing is concrete here. A system assists when it widens the range of voluntary action a person has: it restores a function, removes an obstacle, or offers an option that can be taken or refused. A system controls when it narrows that range, even if the narrowing feels like convenience. The same underlying model can do either, depending on where the objective comes from and whether the person can inspect and veto it. The best future is not the one with the most capable systems. It is the one in which the capability lands on the side of the person’s own aims.
The abundance that widens or concentrates
Every benefit above has a shadow that is the same benefit turned toward power. A tutor that teaches can also hold a learner inside a curriculum someone else designed. A health tool that predicts can also ration, gatekeep, or route people by predicted cost. A forecasting system that protects farmers can also be used to price insurance against them. The technology does not choose the direction; it only makes the direction more efficient.
What decides the direction is ownership, access, and governance. A capability available to everyone tends to diffuse advantage across a population. The same capability held by a few and rented to the rest tends to concentrate it, because the returns accrue to whoever controls the model and the data. This is an institutional and political question rather than a technical one, but it determines whether the positive case arrives at all. Optimism that ignores the distribution is how a good case gets discredited when the benefits fail to spread.
The deepest version of the positive case is therefore not about speed but about dignity. If human worth is not indexed to productivity, intelligence, or enhancement status, then a world of more capable tools can be a world with more room for people to pursue ends that no optimizer would ever score. The purpose of augmentation is to widen participation in meaningful life. Its failure mode is to convert every person into a better-optimized component of somebody else’s system, and to call that progress because the numbers went up.
What would strengthen or break the case
A positive case that cannot be tested is propaganda. This one can be. It gets stronger if capability gains arrive paired with broad access rather than restricted licensing; if independent scientists reproduce the headline results instead of taking press releases at face value; if evaluation and safety practices mature before wide deployment rather than after; and if the institutions that capture the value also distribute it. It gets weaker if the gains stay concentrated, if published claims fail replication, or if systems are deployed at scale before anyone can measure what they do to the people who depend on them.
For scientific claims the standard is the ordinary one: replication, and outcomes that go beyond a benchmark. For moral and spiritual claims the standard is different and should be kept different. Neuroscience can describe the correlates of conviction without adjudicating whether a belief is true. A system can summarize a tradition without becoming the source of its authority. Confusing description with authority is its own failure, and it is not the machine’s fault when people make that mistake; it is ours.
The most useful test is almost embarrassingly simple. Can the person still say no? Can they inspect what the system is doing, choose another, recover their data, and return to the baseline without penalty? A tool that passes is a tool. A tool that fails may still be formally voluntary while being functionally sovereign — pleasant to use, and quietly in charge.
The horizon worth wanting
The best future AI could give us is not a future without effort. It is one in which the expensive parts of being human — understanding a hard thing, regaining a lost ability, learning from someone patient, seeing what a choice will cost before it is made — become cheaper to obtain, and in which the savings flow to more people rather than fewer. That future is not guaranteed by any model, however impressive. It is a destination that has to be chosen deliberately, with the same seriousness we currently apply to building the systems themselves. The capability is arriving on its own. Whether it arrives as a widening of human agency or as a smoother form of dependence is the question the next decade will actually answer.
Sources and further reading
- Jumper et al., “Highly accurate protein structure prediction with AlphaFold,” Nature, 2021
- Royal Swedish Academy of Sciences, “The Nobel Prize in Chemistry 2024,” press release, 2024
- Kestin et al., “AI tutoring outperforms in-class active learning,” Scientific Reports, 2025
- Metzger et al., “A high-performance neuroprosthesis for speech decoding and avatar control,” Nature, 2023
- Lam et al., “Learning skillful medium-range global weather forecasting,” Science, 2023
- Price et al., “Probabilistic weather forecasting with machine learning,” Nature, 2024
- Acemoglu, “The Simple Macroeconomics of AI,” NBER Working Paper 32487, 2024
Loading comments…