Mining Spatial Biology Insights: The Bioinformatician, the AI Scientist, and the Data Integrity…
(Part 3 of the Spatial Biology Series)
Mining Spatial Biology Insights: The Bioinformatician, the AI Scientist, and the Data Integrity Trap
(Part 3 of the Spatial Biology Series)
In **Part 2**, we navigated the hardware trade-offs of the spatial biology universe. We discussed how to choose the right platform by balancing resolution, coverage, and throughput.
But once the assay is run, a new challenge begins. We are left with millions of cells, gigabytes of transcripts, and a coordinate map of tissue architecture. The question shifts from “How do we measure it?” to “How do we mine it?”
Currently, the field is divided into two competing philosophies: the Bioinformatician and the AI Scientist.
While these approaches are not mutually exclusive and often depend on the researcher's technical skill set, they represent fundamentally different ways of seeing the data. More importantly, both are equally vulnerable to the “Achilles’ heel” of spatial biology: data integrity.

Two philosophies for analyzing spatial biology data. Generated by Gemini Nano Banana.
1. The Bioinformatician: The Hypothesis-Driven Approach
The “Bioinformatician” style of analysis is deterministic. Inherited from a decade of single-cell RNA-seq (scRNA-seq) workflows, this approach treats spatial data effectively as a “spatially aware” cell-by-gene matrix.
The workflow is familiar: cluster cells based on gene expression, identify cell types, and then map them back to their X, Y coordinates.
From the perspective of a target validation or mechanism-of-action (MoA) study, the value is obvious. We can visually confirm if a drug target is spatially viable, i.e., whether the target is expressed in an “immune-hot” tumor region or sequestered in an “immune-cold” niche. It offers a grounded, human-interpretable validation of the tissue biology.
The Catch: For discovery, this approach can feel like finding a needle in a haystack. We only find what we already know how to ask for. We search for patterns based on pre-defined rules, which are inherently biased by human intuition.
For example, a common task is to find “cellular neighborhoods” associated with disease states. We might define a neighborhood as “the 30 nearest neighbors” or “cells within 30 microns.” But biology rarely adheres to our rigid parameters. There is no a priori guarantee that the relevant biological interaction occurs at 30 microns rather than 100. By imposing these arbitrary rules, we risk missing the very biology we aim to discover.
2. The AI Scientist: The Data-Driven Approach
In the era of generative AI, the “AI Scientist” offers a seductive alternative: hypothesis-free discovery. Deep neural networks are uniquely suited for this task because they learn latent features directly from the data, without needing humans to define what a “neighborhood” looks like.
We are already seeing the rise of foundation Models, which are large-scale transformers trained on massive datasets to create “digital replicates” of tissue biology. We are also seeing this shift in the industry. In January 2026, Tahoe Therapeutics, the Arc Institute, and Biohub announced joint efforts to build massive atlas datasets comprising over 100 million single-cell perturbation profiles. These large-scale datasets will be used to improve their first-generation virtual cell model. With NIH-funded consortia like the Human Tumor Atlas Network (HTAN) releasing petabytes of spatial data, these sophisticated models are primed to learn the fundamental rules of tissue organization on their own.
The Catch: AI is prone to the “Garbage In, Garbage Out” rule. A foundation model is a statistical sponge; it does not just learn the biology, it also learns the artifacts.
If a biological signal is robust enough to survive bad segmentation or batch effects, it is likely a “gross” anatomical feature we could have seen with simpler tools. The subtle interactions that we bought the expensive machines to find are the first to be drowned out by noise.
The Bottom Line: Data Integrity
Ultimately, data integrity is the common denominator between the two approaches. If the underlying biases in the data are not accounted for, our models will simply hallucinate biological insights, and our spatial analysis will be confounded by imperfect data.
This is not a theoretical risk. A recent study published in Nature Genetics (2026) beautifully demonstrates how segmentation inaccuracies confound downstream analysis.
In a Non-Small Cell Lung Cancer (NSCLC) dataset, researchers compared the gene expression of fibroblasts near the tumor versus those in the distant stroma. The initial analysis was striking: fibroblasts near the tumor appeared to upregulate malignant epithelial markers, such as KRT19. A foundation model might interpret this as a novel “tumor-educated” fibroblast state or a mesenchymal-to-epithelial transition state.
The reality was much more mundane. Researchers found tumor marker genes are often detected in fibroblasts near tumor cells. These tumor transcripts are localized near the cell membrane, where the two cells may have overlapped. It is also possible that the transcripts have “leaked” into the neighboring cell, which is a phenomenon called “transcript leakage”. As a result, cell segmentation algorithms have an almost impossible task to determine the “correct” boundary between two cells, causing the tumor transcripts to be assigned to fibroblasts.

(a) Differential gene expression analysis of fibroblasts near the tumor vs in the stroma. (b) Example fibroblast cells containing tumor markers from the neighboring tumor cell. Tumor genes are often detected near the cell membrane (green) where the two cells may have partially overlapped. Edited from Mitchel et al., Nature Genetics (2026).
Blindly feeding metric tons of such uncorrected data into a foundation model will not solve this problem; it will simply automate the generation of false leads. The model will “learn” that fibroblasts express KRT19, cementing a technical error into the foundation of future discoveries.
Looking Ahead
Before we can build the “Google Maps” of biology, we must ensure our satellite imagery is not blurry. Human-in-the-loop validation remains essential; foundation models won’t catch these errors for us.
In the next article, I will discuss The Boundary Problem, exploring how we can minimize cell segmentation errors in spatial biology data.
Interested in my work? Follow my substack to follow my journey.
메타데이터
- post_id
- 4f12cbb21f74
- slug
- mining-spatial-biology-insights-the-bioinformatician-the-ai-scientist-and-the-data-integrity-4f12cbb21f74
- url
- https://medium.com/@whchou29/mining-spatial-biology-insights-the-bioinformatician-the-ai-scientist-and-the-data-integrity-4f12cbb21f74
- canonical_url
- https://medium.com/@whchou29/mining-spatial-biology-insights-the-bioinformatician-the-ai-scientist-and-the-data-integrity-4f12cbb21f74
- author_url
- https://medium.com/@whchou29
- status
- ok
- fetched_at
- 2026-07-13 06:23:13