AI · Article 59 of 64

AI + Biology + Robotics + Neuroscience: The Feedback Loop That Changes Everything

What happens when the major frontier technologies stop progressing independently and begin accelerating one another?

What happens when the major frontier technologies stop progressing independently and begin accelerating one another?

The intuition behind that question is sound and the usual answer is vague. “Convergence” gets used as a mood, a sense that everything is speeding up at once. The useful version is more specific. A feedback loop exists when the output of one field becomes the input of another, and when the second field’s output improves the first in turn. That is a claim you can check against a paper, and in the last few years several such loops have moved from diagram to bench. What follows separates the loops that have demonstrably closed from the ones that are plausible engineering, and from the ones that remain forecasts, because the three get mixed together constantly and the difference is where all the real questions live.

The shape of a real feedback loop

A loop has three parts: a generator that proposes, an instrument that tests, and a channel that returns the result to the generator so the next proposal is better and the proposal rate rises. When any of the three is missing, you do not have convergence. You have a very good tool being used by a human who is still the loop.

By that standard, most of the excitement of the past decade concerns the first part. Machine learning got dramatically better at proposing things: protein structures, material candidates, molecules, code, robot motions. The interesting news, and the part that is genuinely new, is what has happened to the second and third parts. Robots that can run an experiment unattended, language models that can plan one and operate the hardware, and analysis that closes the cycle without a person in the middle have all been demonstrated publicly in the last three years. That is the difference between better prediction and a closed circuit.

A laboratory that plans, runs, and interprets

The clearest example is the A-Lab at Lawrence Berkeley National Laboratory, reported in Nature in late 2023. The system targets inorganic powders, the kind of material that matters for batteries and catalysts, and it was built to confront a specific asymmetry: computation can propose promising materials far faster than anyone can make them. The A-Lab attacks the making.

Its design is worth stating because it shows exactly which parts of the loop are automated. Recipes come from machine-learning models trained on synthesis procedures text-mined from the literature, on the theory that a new target is best approached by analogy to similar known ones. The candidate targets themselves come from large-scale first-principles databases of phase stability, cross-referenced between the Materials Project and Google DeepMind. Three robotic stations handle powder dosing, heating in a set of furnaces, and X-ray diffraction, with robot arms shuttling samples between them, and everything is reachable through an API so that software agents and humans can both submit jobs. When a recipe fails to hit its target yield, an active-learning algorithm called ARROWS3 proposes a new route using computed reaction energies, and the loop runs again.

The paper reported that the platform synthesized 41 of 58 target compounds over 17 days of continuous operation, a success rate near 71 percent. In January 2026, Nature published an author correction that is more instructive than the original claim. Outside chemists had argued that the diffraction data could not unambiguously identify many of the products, and that some were better explained as known phases. In response, the authors clarified that their novelty claim meant the materials were new to their prediction platform, not necessarily new to science, and that after manually re-analyzing the diffraction patterns the platform had reached the correct conclusion in 36 of its 40 reported successes, with four inconclusive. They also removed one compound that had mistakenly been in the training data.

Both the original result and the correction deserve to be read together, because the correction is not a retraction. The synthesis throughput held up; the identification and the novelty framing did not. The honest summary is that a robot ran a closed loop of propose, synthesize, and characterize on a set of previously unmade materials and got most of them, unattended, in a few weeks, while the harder question of whether the products were correctly understood needed human chemists to settle it.

This is demonstrated science. It is not a general-purpose discovery engine. The target list was drawn from a database of predicted-to-be-stable compounds, which is a strong prior, and the lab worked with air-stable powders, which keeps the chemistry tractable. Extending the same architecture to air-sensitive or organic chemistry is plausible engineering, not a completed result.

When the model drives the bench

The A-Lab closes the loop within one domain, using specialized models. A separate line of work asks whether a general language model can plan and execute an experiment on unfamiliar equipment, which is a harder test because it requires reading documentation, writing code, and recovering from errors without being told how.

That test was run in 2023 and reported in Nature as Coscientist, a system built by Daniil Boiko, Robert MacKnight, Ben Kline, and Gabe Gomes at Carnegie Mellon with collaborators at Emerald Cloud Lab. Coscientist wraps GPT-4 in a set of tools it can call: web search, Python code execution in an isolated container, retrieval over hardware documentation, and an experiment interface that reaches real instruments. Given a plain-language prompt, it can plan a synthesis, look up how a device works, generate the control code, and run it.

Its most-quoted achievement is that it successfully performed palladium-catalysed cross-coupling reactions, the class of chemistry that won the 2010 Nobel Prize and underpins a great deal of pharmaceutical manufacturing. It also controlled a robotic liquid handler to draw simple patterns, used a cloud laboratory’s spectrophotometer to identify the contents of a plate, and optimized reaction conditions by analyzing earlier results. One detail stands out for the convergence argument: when its generated code for a heater-shaker failed, Coscientist consulted the equipment documentation, corrected the code, and retried without human prompting.

That recovery behavior is the part that matters. A system that can only execute a plan it was handed is a tool. A system that can notice its plan failed, find the relevant manual, and fix itself is closer to an agent, and agents are what turn a tool into a node in a loop. The demonstrated scope is still narrow: well-documented reactions, commercial instruments with good APIs, and a supervisor somewhere in the building. The plausible extension is that the same pattern applies to less documented equipment. The speculation is that it generalizes to genuinely novel science, which no result to date establishes.

Biology as a design space rather than a discovery

If the first loop is about throughput, the second is about the design space itself. Biology has always been searched rather than designed, because the mapping from a protein’s sequence to its function is too complex to compute. That is changing, and two results from 2025 show how far.

The first is a generative model for proteins. In Science, the team at EvolutionaryScale, with collaborators at the Arc Institute and UCSF, described ESM3, a model that reasons jointly over protein sequence, structure, and function rather than over sequence alone. Trained with roughly 1e24 floating-point operations and 98 billion parameters on billions of proteins, it is, on the company’s account, among the largest models built for biology. The headline demonstration was a new fluorescent protein, esmGFP, whose sequence is only 58 percent identical to the closest known fluorescent protein. From the observed rate of divergence among natural fluorescent proteins, the authors estimated that generating something this different in one step is equivalent to simulating on the order of 500 million years of evolution.

The phrase “simulating 500 million years of evolution” is doing a lot of rhetorical work, and it should be read precisely. It is not a claim that the model replicated the actual evolutionary lineage of any organism. It is a claim that the model traversed a distance in sequence space that natural diversification would take hundreds of millions of years to cover. That is a genuinely striking result about generative capability, and it is a different thing from having designed a protein that solves an important problem.

The second result is precisely that second thing, and it came from a different group. In Nature in January 2025, Susana Vázquez Torres and colleagues in David Baker’s lab at the University of Washington, working with the Technical University of Denmark and the Liverpool School of Tropical Medicine, used RFdiffusion, a deep-learning structure generator, to design small proteins that bind and neutralize three-finger toxins, the lethal components of cobra and mamba venom. Designing binders for these toxins has resisted conventional antibody production because the toxins provoke a weak immune response in the animals used to make antivenom.

The design pipeline is instructive. RFdiffusion was conditioned to generate protein backbones that would extend a beta-sheet against specific exposed strands on the toxin, a geometric specification rather than a template copy. ProteinMPNN then designed sequences for those backbones, and AlphaFold2 and Rosetta filtered the results before any wet-lab work. After limited experimental screening, the team obtained binders with nanomolar affinity and high thermal stability, and the designed proteins neutralized all three toxin subfamilies in cell assays and protected mice from otherwise lethal neurotoxin doses, with reported survival of 80 to 100 percent depending on the toxin and binder.

This is the loop closing between AI and biology in a form that could matter to people. It is also a good place to be honest about where the result stops. When the researchers moved from purified toxins to whole venom, cell survival fell from 100 percent to somewhere between 70 and 90 percent, and the binders did not reduce tissue damage from a spitting cobra’s venom. No human has received them. Discovery time was compressed dramatically, which is the convergent part; the path from a neutralizing binder to a deployable antivenom is conventional, slow, and mostly outside the loop.

The model grows hands

Everything above still happens in a fume hood or a database. The loop that changes the physical world fastest requires machines that can act in unstructured space, and that capability has been moving.

In March 2025, Google DeepMind introduced Gemini Robotics, a vision-language-action model built on Gemini 2.0 that converts images and language directly into robot motor commands, alongside Gemini Robotics-ER, a companion model specialized for spatial reasoning. The technical report describes a system that generalizes to objects and instructions it was not trained on, executes multi-step manipulation tasks such as folding origami and packing a bag, and adapts to new robot bodies. The team reported that, on its generalization benchmark, the model more than doubled the performance of prior state-of-the-art vision-language-action models. With additional fine-tuning it was specialized to control a humanoid, and later work reported that short-horizon skills could be taught from around a hundred demonstrations.

Read carefully, the result is real and the framing is doing work too. Benchmarks for generalization are constructed by the team that reports them, and “more than doubles” is measured against a moving target. The claims that are firm are that a generalist model can drive several different arm configurations, follow open-vocabulary instructions, and be adapted to a new body without retraining from scratch. The claims that are forecast are everything about reliability in an actual workplace, where a task must succeed thousands of times, not most of the time.

The connection to the rest of the loop is direct and mostly prospective. A robot that can run a chemistry protocol is a way to widen the throat of the A-Lab, letting the same closed loop operate on tasks that today need human hands. That is plausible engineering. Robot labs that discover new science without a human framing the question remain speculation.

Reading and writing the nervous system

The fourth domain is the one where convergence is most often asserted and least often shown. Neuroscience matters to the loop for two distinct reasons: it supplies methods for measuring and intervening in biological systems, and it supplies the brain itself as a target for restoration. The strongest evidence is on the second.

In a 2023 Nature paper, Francis Willett and colleagues reported a speech neuroprosthesis in a participant with amyotrophic lateral sclerosis who could no longer speak intelligibly. Electrode arrays recorded from ventral premotor cortex, a recurrent neural network converted the activity into phoneme probabilities, and a language model resolved those into words. The participant’s attempted speech was decoded at 62 words per minute, roughly three and a half times the previous record and approaching the pace of ordinary conversation, with a 9.1 percent word error rate on a 50-word vocabulary and 23.8 percent on a 125,000-word vocabulary. A related result from a different team, also published in Nature in 2023, reported simultaneous decoding of text, synthesized speech, and facial-avatar animation from a paralyzed participant, at a median of 78 words per minute.

For the convergence argument, the interesting feature is the composition. The result is not one breakthrough but three stacked: better hardware, a neural network decoder, and a large language model doing error correction. Remove any one and the system collapses. That is what convergence looks like at the level of a single product, and it is why progress in one field can make another suddenly practical.

The honest boundary is that this is a research participant with surgically implanted arrays, in a clinical trial, with daily calibration. The demonstrated achievement is restored communication at conversational speed for one person under laboratory conditions. The forecast is portable, low-risk interfaces for people with a wide range of conditions. The gap between them is not mainly algorithmic; it is surgery, longevity of implants, and regulation.

Where the loop accelerates, and where it jams

If the pieces above are real, why hasn’t the promised acceleration arrived all at once? Because loops have bottlenecks, and the bottleneck is rarely the model.

Consider what constrains each stage of a modern discovery cycle. Proposal is cheap and getting cheaper: ESM3 and its kin generate millions of candidates. Filtering is cheap: AlphaFold and Rosetta discard most of them without a pipette. Testing is expensive, slow, and physical, which is why the A-Lab and Coscientist matter more than another model release. And interpretation is again cheap and getting cheaper. In that picture, the whole system’s throughput is set by the second-cheapest thing that is hardest to scale, which is the bench.

That is the most defensible version of the convergence thesis: the loop is real, its weakest link is physical experiment, and the fastest-moving frontier is the one that automates the experiment. It also predicts where the compounding will be uneven. Domains with fast, cheap, well-instrumented assays will accelerate first. Domains requiring living systems, long time horizons, or clinical trials will accelerate last, regardless of how good the models get.

There is corroborating evidence that the general-purpose parts of the loop are themselves improving on a steep curve. METR, a nonprofit evaluation group, measures the length of tasks that frontier AI agents can complete autonomously. In a 2025 study, they found that the task length at which a model succeeds 50 percent of the time has doubled roughly every seven months since 2019, with the strongest early model tested in the original paper, Claude 3.7 Sonnet, sitting near fifty minutes and later models near two hours. The authors are careful that their tasks are software and reasoning tasks, that the trend is a fitted extrapolation, and that performance collapses on messy, underspecified work. Used as an input to the convergence question rather than a prophecy, it says the planning layer of the loop is getting better fast while the physical layer improves more slowly.

What is demonstrated, what is forecast

It is worth stating the tiers plainly, because the convergence topic invites inflation.

Demonstrated, in peer-reviewed or primary sources: a robotic laboratory that synthesized most of a set of predicted inorganic targets unattended over seventeen days, with the identification of its products later refined under outside criticism; a language-model agent that planned and executed named organic reactions, controlled real instruments, and repaired its own failed control code; a generative protein model that produced a fluorescent protein far from any known sequence; a designed-protein therapeutic that neutralized a lethal venom toxin family in cells and mice; a generalist robot model that controls multiple embodiments and follows open-vocabulary instructions; and a speech neuroprosthesis that restored conversational-rate communication to one paralyzed participant.

Plausible engineering, extrapolated from those results: robot labs that handle air-sensitive and organic chemistry; agents that plan and run experiments with weak supervision across a range of equipment; generalist robot models reliable enough for routine laboratory and industrial work; protein therapeutics designed by default rather than by screening.

Speculation, not supported by a result to date: autonomous systems that set their own research agendas and produce conceptual breakthroughs without human framing; a single self-improving loop spanning all four domains; and any timetable for these, since the compounding is bottlenecked at physical experiment and human institutions.

The reason to keep these tiers separate is not pedantry. Each tier supports different conclusions. If a generalist agent can run experiments, then the binding constraint on science becomes lab capacity, and the policy question is how to build that capacity. If the loops only close within narrow domains, the binding constraint is domain expertise, and the question is how to spread it. The two lead to different investments and different fears.

The compression problem

Convergence has a second consequence that has nothing to do with capability and everything to do with time. When the loop between proposal and execution shortens, the interval available for noticing a problem shrinks with it.

The point is easiest to see in the biology case. A system that designs and tests protein binders in weeks rather than years is a gift for treating neglected diseases, and it is also a system that puts functional protein design closer to more people. The same property that democratizes antivenom research makes biosecurity governance harder, because the window in which a dangerous design is visible only in a model, before anyone builds it, gets shorter. Neither the benefit nor the risk is speculative at the level of capability; both follow directly from the demonstrated results above. What is speculative is the severity of the risk, and that is an argument for measurement and screening rather than either complacency or prohibition.

Automating the bench handles the throughput half of that problem and not the judgment half. A robot that runs a thousand syntheses will run the thousand you specified. The A-Lab’s failures were informative because the researchers inspected them and proposed fixes; the system did not decide which failures were scientifically interesting. As the loop speeds up, the ratio of results to people able to interpret them widens, and the scarce resource becomes attention and judgment rather than data. Institutions built for a slower cycle, from peer review to regulatory approval to safety oversight, will feel the mismatch first, and their latency, not the models’, will set the real speed limit for a while.

What to watch

Three questions separate genuine convergence from the ambient hype.

First, where does the loop close without a human in the middle? A model that proposes candidates is not convergence; a robot that tests them and feeds the results back is. Watch the number of experiments a system completes per unit of human attention, and whether that number is rising because of the loop or because of more people watching more screens.

Second, does the automated work survive outside scrutiny? The A-Lab correction is the right kind of event: the claims were checkable, the data were released, and outside chemists found the part that did not hold. The pattern it exposed is general. Automating a measurement is easy; automating the interpretation of a measurement is where the honesty of every downstream claim is decided, and a loop that scores its own outputs will happily optimize against a lenient scorer at machine speed. The healthiest signal in this field is not a spectacular demo but a result that a second lab reproduces using a different stack, which is where the snake-antivenom work is strongest, having been designed, filtered, and then validated in multiple independent assays and laboratories.

Third, does the acceleration show up in outcomes people care about? Cheaper antivenom, faster materials for batteries, restored speech, earlier diagnosis. The loop is worth the disruption only if the compounding lands in the world rather than in benchmark tables. That test is slower to read than any capability curve, and it is the only one that settles the question.

Sources and further reading

Discussion

What would you add or question? Add your comment below. A human reviews it before publication.

Loading comments…

Join the discussion

Comments are public after approval. Please do not include links, email addresses, or private information. For one short AI reply, address @AIGuide in your comment or reply to its opening comment. Cloudflare verifies submissions to limit spam. Read our community guidelines.

The wider community forum is also open: Browse article discussions in the forum · Forum home