The companions to this article ask where the AI race is heading and who is racing. This one asks a narrower and more technical question: why does advanced AI keep arriving in biology? The answer is not primarily marketing. It is structural, and it has to do with the kind of problem biology presents and the kind of instrument a large model is.
Biology is a search problem of almost unfathomable size. A protein’s function is determined by a chain of hundreds of amino acids, each chosen from twenty options; the space of possible chains is larger than the number of atoms in the observable universe. A genome is a string of billions of letters. Yet nature has searched this space efficiently enough to produce every living thing, and evolution’s method is brute force guided by selection over deep time. Humans want to search the same space in years rather than eons. A model that has learned the statistical structure of biological sequences is, in effect, a compressed map of where the good solutions live—and searching a good map is vastly cheaper than searching blind.
That is the core of the argument. Everything that follows is evidence for, or complication of, one claim: that machine learning compresses biological search, and that this compression is why AI organisations keep building toward cells and neurons rather than away from them.
Biology as a compressed map
The clearest demonstration that biological structure is learnable came from protein folding. AlphaFold 3, described in Nature in 2024, predicts the joint three-dimensional structure of complexes containing proteins, nucleic acids, small molecules, ions, and chemically modified residues, using a diffusion-based architecture rather than the residue-pair representations of earlier versions (Abramson et al., 2024). The scientific significance is not that one structure was solved. It is that the relationship between sequence and shape turned out to be learnable at all—that a model trained on known structures could generalise to unseen ones. A learnable relationship is a compressible one, and a compressible relationship is exactly what makes search tractable.
If structure was the first tractable problem, genomes are the second and harder one. In 2026, the Arc Institute and collaborators published Evo 2, a genomic foundation model trained on roughly 8.8 to 9.3 trillion nucleotides drawn from more than 128,000 genomes, with versions up to 40 billion parameters and a context window of one million nucleotides (Brixi et al., 2026). A model that can attend across a million bases is doing something qualitatively different from earlier tools. Rather than scoring one mutation at a time, it can reason about the relationships among distant parts of a genome, which is where much of biology’s meaning lives.
The striking result is not the benchmark score. The authors reported that Evo 2 predicted the functional effect of variants in BRCA1—a gene tied to breast and ovarian cancer risk—with greater than 90 percent accuracy on the zero-shot task. To put that differently: a model that had never been specifically trained to classify BRCA1 variants could nonetheless separate harmful from harmless ones. That is the compression claim made concrete. The model appears to have absorbed enough of what matters about genome function that a clinically relevant question becomes a lookup.
When the model designs rather than predicts
Prediction is useful; design is where the economics change. A map that says where the good solutions live can be inverted to propose new ones.
The Arc Institute team behind Evo demonstrated this directly by generating complete bacteriophage genomes (King et al., 2026). Of nearly 300 designed genomes that were synthesised and tested in the laboratory, sixteen produced viable phages that propagated and inhibited a target strain of E. coli, with the model able to steer which bacterial host the phage would infect. This is a real experimental result, not a computer simulation. It is also a narrow one: bacteriophages are among the simplest organisms known, and the designed genomes were validated against a single bacterial species. Nothing here implies that a model can design a human therapeutic from scratch.
The developers of Evo 2 were careful about the dual-use question, and their decisions are worth recording. They excluded eukaryotic viruses from training, on the grounds that the risk of misuse outweighed the scientific value. They ran red-team exercises asking whether the model could generate pathogenic viral proteins, and reported that the resulting sequences were “effectively random” (Brixi et al., 2026). Those safeguards are the kind of thing a serious lab does. The Perspective published alongside the bacteriophage work argued they are not sufficient on their own, and that claim deserves the same weight as the results.
The loop closes in the laboratory
The most ambitious version of the story is not a better model. It is a closed loop in which a model proposes, an automated laboratory tests, the results feed back, and the model proposes again—compressing the cycle that normally takes a graduate student months into something closer to a script.
A September 2026 result from Anthropic gives a concrete example, and also the honest limits of one. The company reported that Claude, run in an autonomous agent loop for 21.5 hours across 949 sessions and 215.6 million tokens, surveyed reverse transcriptases across roughly 1.9 billion protein clusters and proposed a previously unrecognised genetic system in “jumbo” bacteriophages: a reverse transcriptase, a partner gene, and an array of roughly 200-nucleotide non-coding repeats (Anthropic, 2026). Follow-up bench experiments, performed by human researchers, showed the array is transcribed into distinct short RNAs. The function of the system remains unknown.
Several caveats belong with the result, and they are the difference between a press release and a finding. The preprint has not been peer-reviewed. Rerunning the same campaign ten more times did not reproduce the array discovery, which suggests the result depended on a path the model happened to take. When given the same tools but a reduced budget, the model’s recognition of such systems fell as low as 32 percent. And the authors themselves state that they have not shown the reverse transcriptase is active, that the short RNAs are its substrates, or what the system does. Feng Zhang, whose laboratory specialises in exactly this kind of genome-editing discovery, called the finding “genuinely intriguing”—praise worth having, and also calibrated.
What the result does demonstrate is a change in the texture of scientific labour. The model did not perform the experiment; humans did. What it did was search an enormous space of sequence data fast enough to surface a candidate that a human might never have noticed, then hand it to a laboratory. That is not a discovery machine. It is a very fast research assistant with an unusual tolerance for tedium, and that is a meaningful change in its own right.
The clinical proof, insofar as it exists
For all the excitement about design, the question a sceptic should ask is whether any of it has produced a medicine that helps a patient. So far, exactly one AI-discovered drug has reported human clinical data, and its results define the current ceiling.
Rentosertib, developed by Insilico Medicine for idiopathic pulmonary fibrosis, completed a Phase 2a trial whose results were published in Nature Medicine in June 2025 (Xu et al., 2025). Across 71 patients at 22 sites in China over 12 weeks, the trial met its primary safety endpoint. On the efficacy measure, patients on the 60-milligram once-daily dose showed a mean change in forced vital capacity of +98.4 millilitres, versus −20.3 millilitres for placebo—the treated group’s mean change had a 95 percent confidence interval of roughly 11 to 186 millilitres. That interval describes change within the treated group, rather than the between-group difference.
That is a positive signal, and it is unusually detailed for a first-in-class result. It is also small. The trial enrolled 71 patients in total, spread across dose arms, all at sites in one country, over twelve weeks—too short to establish a durable effect on a progressive disease. Some patients discontinued because of liver toxicity. A single small Phase 2a trial, however well sourced, is not a validation of an entire method; the history of drug development is littered with Phase 2 signals that failed in Phase 3. The defensible claim is narrower: an AI-designed molecule had a tolerability profile supporting further investigation in a small Phase 2a trial and showed a signal worth pursuing. The indefensible claim is that AI has solved drug discovery.
Where design gets specific
The most technically impressive recent results concern antibodies and other proteins the immune system would normally have to evolve. In 2023, the RFdiffusion method showed that a diffusion model could generate protein backbones that fold and function, including binders to specified targets (Watson et al., 2023). Two years later, a follow-up demonstrated de novo antibody design at atomic accuracy: the authors generated single-domain and single-chain antibody variable regions that bound their targets, validated the predicted structures with cryo-electron microscopy, and improved initial binding affinities from the tens-to-hundreds nanomolar range into the single-digit nanomolar range through laboratory affinity maturation (Bennett et al., 2025). Discovering an antibody conventionally means immunising an animal or screening vast libraries. Here a model proposed candidate designs that a laboratory then verified.
Read alongside the phage design and the enzyme discovery, a pattern emerges. Prediction came first, because it is the easiest thing to check. Design came second, because it is the thing that saves money. What has not yet arrived is the third stage: a designed biological product that passes the full gauntlet of clinical trials and reaches a patient. Until that happens, the strongest statement the evidence supports is that models can now propose plausible biological candidates far faster than humans can, and that a meaningful fraction of those proposals survive contact with a laboratory.
The dual-use problem is not hypothetical
Any account of AI in biology that stops at the therapeutic upside is incomplete, and the field’s own researchers are the ones saying so. In the Science Perspective published alongside the bacteriophage work, Thomas Inglesby and Moritz Hanke argued that the ability to compose viral genomes with generative AI now exists while the governance to steer it safely does not, and that the responsibility for preventing misuse should not rest on the discretion of individual labs or on the voluntary screening that synthetic DNA providers currently practise. They called for legally mandated screening of nucleic-acid synthesis orders and of the customers placing them (Inglesby and Hanke, 2026).
Two technical points from that argument deserve attention. The first is that fine-tuning an open model may circumvent training-data exclusions—a safeguard that holds at training time can be undone later, by anyone. The second is subtler. Many biosecurity screening tools work by matching a submitted sequence against a database of known dangerous ones. A genuinely novel design, precisely because it looks unlike anything characterised, may sail past those filters. The better the generative model, the less the existing screening catches, which is an uncomfortable inversion: capability and detectability move in opposite directions.
The underlying result deserves stating plainly, because it is why the biosecurity discussion is not hypothetical: the research team reported the first generative design of complete bacteriophage genomes, confirmed by bench experiments, with some designed phages differing from any known natural phage and one borrowing a DNA-packaging protein from an evolutionarily distant virus. This is not an argument against the research. It is an argument that the research has become powerful enough that its externalities need managing by institutions rather than by good intentions.
Why the migration keeps happening
Pull the threads together and the migration makes sense for reasons that have nothing to do with fashion.
Biology is a domain where the data is abundant, the ground truth is expensive but obtainable, and the underlying regularities are learnable. Those are precisely the conditions under which a large model outperforms a hand-built tool. The bottleneck in biology is not ideas; it is experiments, which cost money and time and can only be run so fast. A model that narrows the space of experiments worth running is therefore valuable in direct proportion to how well it predicts, not how impressively it writes.
The laboratory loop amplifies this. When a model proposes and a robot tests, the loop can run at a speed no human team can match, and each cycle produces data that improves the next proposal. AI organisations that want to keep improving their models also want more of the highest-quality data available, and a laboratory they control is a source of exactly that. The interest in biology and the interest in better models reinforce each other. The move into biology extends the AI business rather than diverting from it.
The convergence is also visible from the neuroscience side, where the same logic applies to signals rather than sequences. Neural recordings are noisy, high-dimensional, and full of structure that classical decoders struggle to extract. Large models are unusually good at exactly this kind of decoding, which is why they appear in brain-computer interface research alongside the electrodes and the surgery. The subject differs—neurons rather than nucleotides—but the mechanism of advantage is identical: abundant data, learnable structure, a bottleneck elsewhere that better prediction can relieve.
What would change the assessment
The claim of this article is that AI is moving into biology because models compress biological search, and that the compression is real but partial. Two kinds of evidence would sharpen it in either direction.
Evidence that it is working would look like a designed biological product—a drug, an antibody, a gene therapy—reaching patients through full clinical trials, not just a Phase 2a signal. It would look like laboratory loops reporting results that independent groups reproduce, rather than single runs whose outcome depends on the path a model happens to take. And it would look like governance that demonstrably keeps pace, so that the screening systems protecting against misuse are not defeated by the very novelty the models produce.
Evidence that it is overrated would look like Phase 2 failures that mirror the base rate for conventionally discovered drugs, which would suggest the compression is not translating into clinical advantage. It would look like irreproducible discovery claims that turn out to be artefacts of a single campaign. And it would look like a regulatory environment that treats computational design as a compliance formality rather than a shift in capability.
None of that evidence is in yet. The state of the field is a set of real, carefully reported results—a structure predictor, a genome model that designs working phages, an antibody designer validated by cryo-EM, an autonomous agent that surfaced an unexplained genetic system, and one small clinical trial—surrounded by a much larger volume of claims that have not been tested. Whether the migration into biology transforms medicine or settles into a well-funded research program depends on which of those two piles grows faster. The honest answer, today, is that we have the first promising pieces and not yet the pattern.
Sources and further reading
- Josh Abramson et al., “Accurate structure prediction of biomolecular interactions with AlphaFold 3,” Nature, 2024
- Garyk Brixi et al., “Genome modelling and design across all domains of life with Evo 2,” Nature, 2026
- Anthropic, “Claude discovers a novel enzyme system with CRISPR-like repeats,” 2026
- Zhihong Xu et al., “A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial,” Nature Medicine, 2025
- Joseph Watson et al., “De novo design of protein structure and function with RFdiffusion,” Nature, 2023
- Nathaniel Bennett et al., “Atomically accurate de novo design of antibodies with RFdiffusion,” Nature, 2025
- Megan King et al., “Generative design of bacteriophages with genome language models,” Science, 2026
- Thomas V. Inglesby and Moritz S. Hanke, “AI-designed viral genomes,” Science, 2026
Loading comments…