ESMFold2: The Bio-AI Game Is Elimination Now
Biologics ideation is becoming table stakes. The real game is now high-throughput filtering, trustworthy liability prediction, and turning…
ESMFold2: The Bio-AI Game Is Elimination Now
Biologics ideation is becoming table stakes. The real game is now high-throughput filtering, trustworthy liability prediction, and turning binders into drugs. ESMFold2 is impressive confirmation of that shift, not the end of it.

Biologics ideation is now table stakes. The core challenge of drug development is surviving the elimination gauntlet.
This is the fourth in a four-part series that formed a dialog I didn’t entirely plan going in (but where an unexpectedly rich stream of recent AI-related releases pointed me). The Breakfast Test laid the foundation — what it actually means to evaluate an AI release as science. The Reproducibility Tax proposed the MVD disclosure standard: the floor below which a corporate announcement stops being science and becomes marketing. Last week’s essay introduced Rocklin’s MGnify stability dataset, 1.8 million empirically measured folding stabilities that point to what generative models are actually missing. This week, Biohub’s ESMFold2 closes the loop — and, not incidentally, is what MVD looks like when someone actually meets it. When I return, I’ll re-widen the aperture — because while AI is interesting, it’s not everything.
I’ve spent years thinking about the biologics drug space. It’s not where I started, but I wound up in the middle of it when Schrödinger asked me to put together the first commercial biologics-focused package — BioLuminate. At the time, novel biologics platforms came slowly, and de novo ideation was limited to a small number of labs, like David Baker’s (who subsequently won the Nobel Prize for this work). You’d attend a conference and Baker would present a deck I’d wryly call the “Bob the Protein Builder deck,” alternating “Here’s the target, can we build it?” and “Yes we can.” You’d watch in amazement, knowing it wasn’t that simple except when distilled to the successes — each of which represented years of hard work and deep expertise. Still, it was amazing work, full stop. Ideation was the limit. And even if many paths were more dead end than new gate, you could see where the field was going, albeit slowly.
In the past few years, post-AlphaFold, ideation has gone from a bespoke enterprise of amazing surprises to something we can envision at scale, to something actually being delivered at scale. ESMFold2 is the latest payoff in that story.
Still, I’ve written about enough of these moments to recognize the pattern before the press kit arrives. Protein-systems hype cycles in the AI era always trigger the same two reflexes. One is the eye-roll: same story, bigger model, same limitations, better branding. The other is the opposite mistake: a breathless claim that biology has been solved, evolution outpaced, and the machine transformed into an all-seeing oracle. I’ve watched this field long enough to know both reflexes are wrong in the same way — they let you stop thinking.
In late May, Biohub announced ESMFold2 (in an impressively detailed 106-page publication that should serve as a model for AI model disclosure). This is the field’s latest, and arguably most impressive, step toward routine biologics ideation, and a direct confirmation of the argument I made when BoltzGen arrived last November: that the bottleneck had shifted from ideation to filtering and liability. The ideation capability has taken another genuine leap forward. The liability gap has not.
ESMFold2 is not just AlphaFold with extra parameters and a sharper press kit. And it is certainly not a physics engine in disguise. What it is, instead, is a system that has learned enough of protein biology’s statistical grammar to do serious work without being handed the usual evolutionary scaffolding at inference time.
That is the part worth paying attention to. It shifts the center of gravity from alignment to inference. But it does not shift the center of gravity from binder to drug. And that distinction matters more than almost anything else in this field.
Unexpectedly impressive and game-changing are not necessarily the same thing.
What Actually Changed
For a long time, structure prediction was built around the MSA worldview. If you wanted the model to fold a protein, you gave it the protein’s relatives — the evolutionary gossip, the coevolutionary breadcrumbs. If residue A and residue B had been mutating together for millions of years, that was probably not a coincidence. The power of cross-correlation to imply spatial proximity is an idea that allowed Rosetta to redefine what was possible. AlphaFold turbocharged it into something that changed the field entirely.
This worked extremely well when nature had left a clean trail of homologs behind. It worked far less well for orphan sequences, de novo designs, and proteins without a convenient evolutionary entourage.
ESMFold2 changes the workflow in a way that matters more than it first sounds. It predicts structure from a single sequence, without requiring an explicit MSA at inference time. That does not mean evolution has vanished; it has been absorbed. The model was trained on approximately 2.8 billion sequences — roughly two orders of magnitude more than its predecessor ESM2 — spanning bacteria in deep soil, extremophiles, and metagenomic sequences from ocean and earth environments that have never been fully characterized. What it learned is a compressed representation of the rules that make protein space hang together: not magic, but compression with real boundaries.
There was no guarantee this would work. The further you push between the raw encoding — sequence alone — and the structural readout, the more you’re asking the model to have internalized something that previously required explicit evolutionary annotation. This project assuredly started with the question “Can this possibly work?” rather than “This should work.”
It worked.

ESMFold2 bypasses explicit multiple sequence alignment, predicting protein structure directly from a single sequence at inference time.
Why That Matters
The strongest argument that this is not just another scaling story is not the benchmark deck. It is the lab.
Designing a binder from scratch is precisely where “looks plausible on screen” usually collides with reality. Biohub designed de novo minibinders against five disease-relevant targets — EGFR, PDGFRβ, PD-L1, CTLA-4, and CD45 — using ESMFold2, and those designs worked. The most striking result: PD-L1 minibinders and scFvs restored T-cell signaling in a functional assay with IC50 values of 1.6 nM and 39 nM, respectively — comparable to 2.6 nM for an atezolizumab-derived control. Atezolizumab is an FDA-approved checkpoint therapy. The designed binders were competing with it. That is where rhetoric stops and molecules start.
The structural validation is equally pointed. Cryo-EM confirmation of an EGFR minibinder showed the experimentally determined complex aligning with the predicted structure at 1.204 Å RMSD. The model predicted the binding pose. The electron density confirmed it. That is not a benchmark number — it is a real structure.
The interpretability angle is more than a side note. Sparse autoencoder work suggests the model has internalized recurring biological motifs, including the nucleophilic elbow — a specific catalytic feature that a single learned representation activates across structurally unrelated enzymes using entirely different scaffolds. That does not mean the model “understands” proteins the way a biochemist does. But a representation that independently recovers that motif, unsupervised, is pattern recognition with real anchors, not statistical theater.
One phrase in the above deserves parsing: “those designs worked” means they bound their targets. It does not mean they are drugs. That gap is where this story gets harder.
Where the Story Breaks
This is where the hype narrative starts to wobble.
ESMFold2 is not a physics engine in disguise. It is a powerful statistical model trained on an enormous slice of evolutionary life, which means it remains bounded by what that slice contains.
De novo designs. The farther you drift from familiar sequence space, the more the model reveals what it is: an interpolator. Push it far enough out of distribution and hallucinations return. The paper itself is careful here: the validated designs emerged from large-scale search, not single-shot generation. Hit rates of 36–88% for minibinders and 15–29% for scFvs are impressive. No, scratch that. They are astonishing hit rates for de novo design. But they are still hit rates.
Conformational diversity. The model also inherits the familiar single-structure bias. Biology loves ensembles, alternates, context dependence, ligand effects, mutations — all the mess that makes proteins interesting. The model, like most structure predictors (and most structural biologists!), wants a single answer. Sometimes that answer is useful. Sometimes it is a tidy lie. The paper’s own characterization of designs as covering diverse binding orientations and structural topologies is encouraging, but it does not address what the model cannot see: the dynamic ensemble around those structures.
The PTM blind spot. Often ignored but tremendously relevant to biologics drugs are post-translational modifications. Scientists frequently minimize them because they are hard. Nature never does. PTMs represent a level of predictive specificity that is beyond most physics-based methods and entirely beyond current AI capabilities — which means every model in this space, however impressive, is working with a map that excludes a substantial fraction of the biology it claims to predict.
Isoforms and edge cases. Isoforms, splice variants, and similar biological irregularities are another reminder that protein biology refuses to be a clean sequence-to-structure prompt. These are the cases where statistical fluency starts to fray, because the biology has stopped behaving like the training set.
Biosecurity. The paper handles this carefully; most coverage skips it. A model this capable of navigating protein function space across taxonomic neighborhoods has dual-use implications that are not hypothetical. The open license compounds this: a closed API can be monitored; a downloaded model cannot. Models trained to understand functional motifs across sequence space can, in principle, aid the design of proteins with harmful biological activity — including toxins and pathogen components — with less domain expertise than was previously required. This is not a reason to dismiss the work. It is a reason to name it plainly, because the field’s tendency is to treat it as a footnote until it isn’t.
The Liability Gap
The most important limitation has nothing to do with the model’s architecture.
Even when ESMFold2 designs a binder that works — genuinely, experimentally works — it has solved only the first step. In biologics, the distance between a functional hit and a developable drug is where most programs die. De novo scaffolds are non-self by definition. Immunogenicity risk is baked in from the start. Designed minibinders have no track record of PK stability; most would be proteolyzed in minutes in vivo. Expression yield, aggregation behavior, and manufacturability are nowhere in the model’s scoring function. The paper acknowledges this directly — the designs were characterized for homogeneity and non-aggregation by size-exclusion chromatography, which is a start, but it is a long way from an IND package.
None of this is ESMFold2’s fault. It was not built to solve those problems. But it is worth saying clearly: the model has dramatically accelerated ideation while leaving the liability landscape exactly where it was.
Rocklin’s MGnify database of 1.8 million empirically measured, publicly accessible, folding stabilities that I wrote about last week is exactly the class of infrastructure the liability side of this problem actually needs — not more generative capability, but more empirical boundary. Wet-lab reality rather than computational proxy. Those ΔG values will likely be more valuable in five years than they are right now, because the models capable of fully exploiting them haven’t been built yet. The Biohub work is, in real time, demonstrating exactly why.

Generative models yield computational proxies. Closing the liability gap requires anchoring those predictions in empirical wet-lab reality.
This isn’t just a theoretical critique; it is a practical wall that the generative community is hitting in real time. Look no further than the technical report for BoltzProt-1 that dropped this week. The authors didn’t announce a radical leap in generative physics. Instead, they spent much of their paper detailing a brutal two-stage gauntlet: a novel interaction-scoring model, BoltzPPI, trained on the “museum of winners” in the PDB and patent databases to filter candidates before experimental testing, followed by a comprehensive developability panel measuring thermal stability, hydrophobicity, polyspecificity, and monomer purity via analytical size-exclusion chromatography. The generative engine still can’t design against aggregation upstream — it has to be screened out after the fact. When a flagship structural framework celebrates a 58% survival rate across physical liability filters as a headline result, that’s a pretty good don’t-just-trust-me demonstration that the field is catching on to where the real problem lives.
Where This Leaves Us
We are closer than we have ever been to the point where identifying a biologic target and generating something that binds it well is a tractable, fast, and reliable problem. That is real, and it deserves to be said without hedging. But the liability landscape has not moved at anything close to the same pace. In many cases, it hasn’t moved at all.
We can generate binders with growing confidence. We still cannot predict, with any reliability, whether those binders will express cleanly, survive in vivo, avoid triggering an immune response, avoid aggregation, retain suitable viscosity at point of injection, or stay away from the ten thousand things we did not intend them to touch.
The larger commercial prize will belong to organizations that can close that gap — by automating liability assessment at scale, by building the datasets that make computational liability prediction trustworthy, or both. The companies that figure out how to generate clean, systematic data on immunogenicity, PK, aggregation, and specificity — and build models trained on that reality rather than on structural proxies — are the ones that will turn this ideation revolution into a development revolution.
ESMFold2 demonstrates how much of protein biology’s statistical grammar can be compressed into a learned representation. That achievement is real, arguably shockingly good, and it matters. But drug development is ultimately a game of elimination, not generation. The industry’s central challenge in biologics is no longer imagining candidate molecules. It is discovering which of them can survive the gauntlet that begins after generation ends.
ESMFold2 is not an oracle, and it is not a simulator. It is a remarkably fluent model of protein biology. The hardest questions remain. That is not a failure of the model — it is a description of where the frontier has moved.
I write A Half Life on Substack — observations from drug discovery, computational chemistry, AI, and how the field actually works. Sometimes just life. Free to follow.
메타데이터
- post_id
- bbf2db78d6a5
- slug
- esmfold2-the-bio-ai-game-is-elimination-now-bbf2db78d6a5
- url
- https://medium.com/@dapscience/esmfold2-the-bio-ai-game-is-elimination-now-bbf2db78d6a5
- canonical_url
- https://medium.com/@dapscience/esmfold2-the-bio-ai-game-is-elimination-now-bbf2db78d6a5
- author_url
- https://medium.com/@dapscience
- status
- ok
- fetched_at
- 2026-06-20 20:29:01