The Next Biomedical Revolution Won’t Come From AI — It Will Come From This Missing Data
The current wave of biomedical AI hype centers on models: bigger transformers, better structure predictors, more capable generative…
The Next Biomedical Revolution Won’t Come From AI — It Will Come From This Missing Data
The current wave of biomedical AI hype centers on models: bigger transformers, better structure predictors, more capable generative chemistry. That focus is understandable given how visible model breakthroughs are. But a strong case can be made that the field’s next real leap forward is not waiting on a better algorithm — it’s waiting on data that doesn’t fully exist yet at the scale, resolution, and integration needed. Better models trained on incomplete data are still fundamentally limited by what that data can represent, and three specific gaps stand out.
Single-cell resolution: what bulk data has always hidden
For most of genomics’ history, sequencing measured a tissue sample as one averaged signal — thousands or millions of cells blended together. That average conceals enormous biological reality: a tumor isn’t one population, it’s a heterogeneous mix of cell states with different mutations, different drug sensitivities, and different roles in progression, and a bulk measurement reports something closer to their average than the truth of any individual cell. Single-cell sequencing technologies have started correcting this by measuring individual cells rather than populations, revealing rare cell subpopulations, cell-state transitions, and heterogeneity that bulk data structurally cannot detect no matter how sophisticated the downstream model is.
This matters directly for drug development and disease modeling: a treatment that looks effective against an averaged bulk signature can still fail against a resistant minority subpopulation invisible in that average — a pattern increasingly recognized as a driver of treatment resistance and relapse.
Spatial biology: where a molecule is matters as much as whether it’s there
Standard sequencing, single-cell or bulk, requires dissociating tissue — breaking it apart into a soup of cells or molecules for measurement. That process destroys spatial information: which cells were physically adjacent, which cell sat within a particular tissue microenvironment, how a signal gradient was organized across a structure. Spatial biology methods — spatial transcriptomics and spatial proteomics — preserve this positional context, measuring gene or protein expression while keeping track of where in the tissue that measurement came from.
This turns out to matter enormously for understanding disease. Immune cell infiltration patterns in a tumor, the specific architecture of a fibrotic lesion, or the layered organization of a diseased tissue all carry information that a dissociated, position-blind measurement simply cannot capture, no matter how many cells it profiles.
Multi-omics: no single layer tells the whole story
Genomic, transcriptomic, proteomic, epigenomic, and metabolomic data each capture a different layer of biology, and each layer alone is an incomplete picture. A gene can be present and structurally normal (genomics) while being transcriptionally silent (transcriptomics), or transcribed but not translated efficiently (proteomics), or translated into a protein that’s chemically modified into an inactive state (epigenomics and post-translational data). Predicting a phenotype from any single layer discards information the other layers hold, and integrating multiple omic layers on the same samples — rather than analyzing them in separate studies on separate cohorts — is technically and logistically difficult, which is exactly why comprehensive, well-integrated multi-omic datasets remain comparatively rare.
Why this is a data problem, not (only) a modeling problem
A model — however architecturally sophisticated — can only learn patterns present in its training data. Feeding a state-of-the-art model bulk, dissociated, single-layer data means it’s fundamentally blind to cell heterogeneity, spatial organization, and cross-layer relationships, no matter how much compute or how many parameters are thrown at the problem. This is a meaningful part of why some of the field’s most-hyped predictive tools still show a persistent gap between benchmark performance and real clinical or experimental usefulness — they were trained on data that already discarded information the underlying biology actually depends on.
What closing the gap would look like
The infrastructure is being built, gradually: single-cell atlases profiling cell types and states across tissues and diseases, spatial datasets mapping molecular signals onto tissue architecture, and multi-omic cohorts pairing several data layers on the same patients over time. These efforts are expensive, technically demanding, and slower to produce than a new model architecture — which is exactly why they’ve received comparatively less attention despite arguably being the more binding constraint.
The next real jump in what biomedical AI can actually predict may have less to do with a smarter algorithm and more to do with finally giving existing algorithms the resolution, spatial context, and cross-layer integration that biology has always required and that most current datasets still don’t provide.
Zeemal Fatima Nawaz is an independent computational biology researcher and first author of a preprint on BEST1/Bestrophin-1 variant pathogenicity (DOI: 10.21203/rs.3.rs-9389974/v1).
메타데이터
- post_id
- 4d2aed1839c2
- slug
- the-next-biomedical-revolution-wont-come-from-ai-it-will-come-from-this-missing-data-4d2aed1839c2
- url
- https://medium.com/@zeemalfatima078/the-next-biomedical-revolution-wont-come-from-ai-it-will-come-from-this-missing-data-4d2aed1839c2
- canonical_url
- https://medium.com/@zeemalfatima078/the-next-biomedical-revolution-wont-come-from-ai-it-will-come-from-this-missing-data-4d2aed1839c2
- author_url
- https://medium.com/@zeemalfatima078
- status
- ok
- fetched_at
- 2026-08-06 11:15:05