In May 2025, Google DeepMind described a coding agent, AlphaEvolve, that had been making small but real changes to the company’s own computing stack. It found a scheduling heuristic for placing jobs across data centers that has been in production for over a year and continuously recovers, on average, about 0.7 percent of Google’s worldwide compute. It proposed a rewrite of a Verilog arithmetic circuit inside a Tensor Processing Unit that removed unnecessary bits while preserving the circuit’s function, and that change was verified and folded into a forthcoming chip. It reworked a tiling heuristic for a matrix-multiplication kernel used in Gemini’s training, speeding that kernel by 23 percent and trimming total training time by about one percent. It also discovered a way to multiply two 4×4 complex matrices using 48 scalar multiplications — the first improvement over Strassen’s 1969 method in that setting, published, provably correct, and bordered by the knowledge that nobody has shown 48 is optimal.
That list is worth staring at because it looks nothing like the pictures most people carry around when they hear “intelligence explosion.” There is no single mind racing away from its makers in a sealed room. There is a tool making the machinery that trains AI somewhat cheaper, and the savings being reinvested into training more AI. It is recursive, but it is recursive through supply chains, chip layouts, compilers, and job schedulers — through hardware and energy and the speed of a feedback loop measured in weeks rather than nanoseconds.
The question this article takes up is therefore narrower than “will AI explode?” and, I think, more useful. Does an intelligence explosion require a sudden, discontinuous jump in capability, or can something that deserves the name arrive gradually, as an accumulation of compounding improvements in research and engineering? The answer depends on separating three things that get braided together in most discussions: what has been demonstrated, what engineers could plausibly build next, and what remains speculation dressed as forecast.
What the term originally meant, and what it quietly imports
The phrase “intelligence explosion” comes from the mathematician I. J. Good, who wrote in 1965 that an ultraintelligent machine would be the last invention humans need make, because such a machine could design better machines, and the sequence would run away. The idea migrated into contemporary discussion largely through Nick Bostrom’s 2014 book Superintelligence, which distinguished a “fast takeoff,” in which a system’s capabilities surge past human control within a short span, from a “slow takeoff,” in which the same destination is reached over years or decades.
That vocabulary imprints a picture. A takeoff is a physical event with a moment to it — the instant the wheels leave the ground. Much of the public argument about AI timelines is really an argument about that moment: whether it comes, and when, and how quickly afterward the world changes. But a system can be explosive in the sense that its growth rate rises without bound and still take a decade to pass through the transitions that matter. Compounding is a different shape from suddenness, and the two are being used to argue about the same thing.
The distinction is not academic. If progress is fundamentally continuous — each year’s models a little more capable, each year’s infrastructure a little more efficient — then the levers that matter are institutional: how fast discoveries diffuse, who can inspect and veto deployed systems, whether verification keeps pace with capability. If progress can be genuinely discontinuous, the levers are different and much more fragile, because there may be no time to pull them. The evidence available today supports the first picture strongly, the second weakly, and the honest task is to say which observations would move the estimate.
The most measurable trend in the field
The clearest quantitative result in this area comes from Model Evaluation and Threat Research, or METR. In a March 2025 paper, the group proposed a way to compare models not by benchmark scores but by the length of the tasks they can complete. They assembled 170 software and research-engineering tasks, timed how long skilled humans took to do the same tasks, and then fit a curve that predicts a model’s success rate as a function of that human time. The task duration at which a model succeeds half the time became its “50 percent time horizon.”
The striking finding is that this horizon has been growing exponentially, roughly doubling every seven months from 2019 through 2025. GPT-2 in 2019 sat near tasks that take a person two seconds; by early 2025, Claude 3.7 Sonnet sat near tasks that take a person about fifty minutes (Kwa et al., 2025).
Several features of this result matter more than the headline number. First, the increase appears to be driven less by raw knowledge than by reliability — models stringing together longer sequences of actions without losing the thread. Second, the group is candid about limits: the tasks are software and reasoning tasks that can be automatically scored, which is exactly the domain where training works best; performance drops on “messier” tasks with ambiguous goals and slow feedback; and the whole measurement is a comparison against particular human baseliners, so the absolute numbers carry real uncertainty even if the trend is robust. Third, the trend may or may not continue. The authors’ own extrapolation predicts month-long autonomous software tasks within a few years if it holds — and the sentence “if it holds” is doing most of the work.
A 2025 follow-up by the Oxford economist Toby Ord offered a mechanistic reading of the same data. If an agent fails at a roughly constant rate per unit of task time, success decays exponentially with task length, and each model gets a kind of “half-life” measured in human-hours (Ord, 2025). That model fits METR’s data well, and it suggests why horizons rise as models get more reliable: a task that is a thousand small steps is only completed if essentially every step holds. The same arithmetic implies that very high reliability is much harder to reach than a median-level horizon, and that the gaps between “half the time” and “almost always” are wide.
Where the loop actually spins
If there is a feedback loop, the interesting question is where it bites. AlphaEvolve is a good specimen because its results are specific and checkable. It improved a data-center scheduler, a chip circuit, and a training kernel — each a place where a small percentage gain, multiplied across an enormous and repeated workload, becomes consequential. It also found genuinely new mathematical constructions, rediscovering the best known solutions on roughly three-quarters of more than fifty problems and improving on the best known result in about a fifth of them, including a new lower bound for the kissing-number problem in eleven dimensions (Novikov et al., 2025).
Read the method and the boundary appear together. AlphaEvolve needs an automatic evaluator: a program that can score a candidate solution without a human in the loop. In the domains where that exists — mathematics, algorithms, kernel performance, circuit correctness — it can search far more broadly than a person. In domains where judgment about a physical or social world cannot be reduced to a function, it is out of scope. That is a demonstrated engineering fact, and it locates the current feedback loop precisely: it is strong where success is machine-checkable and weak where it is not.
The complementary demonstrated fact comes from RE-Bench, METR’s benchmark of realistic machine-learning research-engineering tasks. There, the best AI agents beat human experts when both were given two hours, but humans improved more as the time budget grew, and at thirty-two total hours the best humans roughly doubled the top agent’s score (Wijk et al., 2024). Agents were far cheaper per unit time and generated candidate solutions more than ten times faster, yet the median agent attempt barely improved on the starting solution. Fast iteration and unreliable judgment coexist. That pattern — cheap breadth, fragile depth — is the shape a compounding takeoff would have to overcome, and it is not obvious that scaling compute brute-forces it.
Explosive in the economists’ sense
The word “explosion” in the economic literature means something precise: growth in world output at an order of magnitude above today’s roughly three percent per year. Ege Erdil and Tamay Besiroglu reviewed the arguments for and against it and concluded that growth theory, taken at face value, predicts accelerations once AI can substitute for labor across most tasks — because labor is the one key input that has not been accumulable at will, and an artificial workforce removes that bottleneck (Erdil and Besiroglu, 2023). Their own verdict was deliberately hedged: the odds of explosive growth this century following widespread automation are about even, with high confidence in either direction unwarranted.
A more recent and more pointed result comes from Tom Davidson, Basil Halperin, Thomas Houlden, and Anton Korinek, who modeled research as a network in which progress in one sector raises productivity in others. They derive a clean condition under which growth becomes superexponential: it happens when the combined strength of a technological feedback loop and an economic feedback loop — better tools enabling more research, and more output financing more research — overcomes the diminishing returns that ordinarily make ideas harder to find. In a simulation calibrated to observed trends, fully automating software research with only modest automation elsewhere produced a singularity within six years (Davidson et al., NBER, 2026). I take that figure as a demonstration of a mechanism, not a forecast. It shows that a compounding loop can, under stated assumptions, become explosive in the mathematical sense without any single magical jump of the kind the popular phrase conjures.
Epoch AI’s GATE model reaches a similar qualitative place by a different route. Feeding scaling laws and growth theory into one framework, it finds large investment in AI compute could be justified and that growth accelerates markedly once a large fraction of tasks is automated — but its authors stress that the model assumes embodied AI can substitute for human labor across essentially all tasks, and that it omits robotics, diffusion, and market frictions (Epoch AI, 2025). The result is a picture of accelerating growth emerging over roughly a decade as automation spreads, not a step change.
The brake nobody can remove, and the one nobody is sure about
Two things stand between a well-specified feedback loop and an actual explosion. The first is the difficulty of discovery itself. Epoch’s estimate of the “returns to research effort” in software — the parameter that decides whether an automated research workforce produces more-than-proportional or less-than-proportional gains — came out around 0.83 for the chess engine Stockfish, with a standard error of 0.15 (Besiroglu, Erdil, and Ho, 2024). That value sits just below the threshold of 1.0 that separates accelerating from decelerating progress. It does not rule out a crossing, and it says nothing definitive about AI research specifically, but it is the closest thing to a direct measurement and it leans against the runaway version.
The second brake is physical and institutional. Even a software feedback loop runs on chips, power, and data, and the gains AlphaEvolve found are exactly the kind that compound only because they are applied across an enormous, repeated workload. A loop that must buy fabs and build power plants moves at construction speed. Regulation, export controls, and the concentration of capability in a handful of firms can damp the loop too, though Erdil and Besiroglu argue these frictions are unlikely to halt it indefinitely given the incentives at stake. Which of these binds first, and when, is genuinely open.
What would actually distinguish the two pictures
If progress is compounding, we should expect a consistent pattern rather than a break: horizons that lengthen smoothly, automation that spreads task by task and checkable evaluator by checkable evaluator, and gains that show up first where verification is cheap. That is what the record shows so far. If it is explosive in the sudden sense, we should eventually see the trend steepen rather than merely continue — frontier systems improving faster than the exponential line, not along it — and, crucially, the returns to research effort climbing past 1.0 rather than hovering just below. Both statements are falsifiable, which is what makes them worth writing down.
There is also a subtler discriminator. In the quick version of the story, capability outruns everything else and the transition is over before institutions respond. In the compounding version, capability and its consequences diffuse at different rates, and the binding constraint becomes how long it takes a discovery to reach the laboratory, the clinic, the factory, and the law. The RE-Bench result — that agents are fast and cheap while humans pull ahead given time — is a hint about which picture is truer today. So is the fact that AlphaEvolve’s largest wins were percentage improvements to infrastructure that required years of accumulated data-center and compiler work to even be worth optimizing.
None of this settles timelines, and it should not be read as doing so. The honest statement is that a compounding, infrastructure-mediated feedback loop is demonstrated at small scale; that economic models show such a loop could become explosive under conditions that are plausible but unverified; and that a sudden, discontinuous takeoff remains speculation — possible, argued from mechanism, and unsupported by anything measured. The three claims are ranked by evidence, and the ranking matters.
The consequence for how to think about it
If the compounding picture is closer to right, then the central risk is not that a single system awakens and outpaces the world in a weekend. It is that a system of already-accelerating loops outruns the slower human processes of verification, consent, and accountability — while remaining, in each individual step, unremarkable. A scheduling heuristic that recovers 0.7 percent of compute, a circuit rewrite, a training kernel that is a fifth faster: none of these is dramatic, and all of them feed the next round. The danger of the gradual version is precisely its legibility as ordinary engineering.
That is why the useful posture is neither to dismiss the explosion framing nor to adopt its most cinematic form. It is to watch the specific joints where a compounding loop could become a discontinuous one, and to insist that the machinery of oversight — audit, portability, the ability to inspect what a system is doing — scale at something like the rate of the loops it is meant to govern. An intelligence explosion that arrives as a decade of compounding has one great advantage over the sudden kind: it leaves time to act, if anyone chooses to.
Sources and further reading
- Kwa et al., “Measuring AI Ability to Complete Long Tasks,” arXiv, METR, 2025
- Wijk et al., “RE-Bench: Evaluating Frontier AI R&D Capabilities of Language Model Agents against Human Experts,” ICML / METR, 2024
- Novikov et al., “AlphaEvolve: A Coding Agent for Scientific and Algorithmic Discovery,” Google DeepMind, 2025
- Ord, “Is There a Half-Life for the Success Rates of AI Agents?,” arXiv, 2025
- Erdil and Besiroglu, “Explosive Growth from AI Automation: A Review of the Arguments,” Epoch AI, 2023
- Davidson, Halperin, Houlden, and Korinek, “When Does Automating AI Research Produce Explosive Growth?,” NBER Working Paper 35155, 2026
- Epoch AI, “GATE: Modeling the Trajectory of AI and Automation,” 2025
- Besiroglu, Erdil, and Ho, “Do the Returns to Software R&D Point Towards a Singularity?,” Epoch AI, 2024
Loading comments…