Long-read (Nanopore) vs Short-read (Illumina) Data in Cancer Genomics
Why do we sequence the human genome? Imagine for a second that you don’t how to read and someone asks you to compare 2 books…
Long-read (Nanopore) vs Short-read (Illumina) Data in Cancer Genomics
Imagine for a second that you were never taught how to read. Someone walks up to you and hands over two heavy, gorgeous books, say 3000 pages long, and tells you to find everything that’s different. Now, you’re smart; just unlettered. There are so many ways you might approach that problem. You could first check if they even have the same pages. You could compare the designs and say which is prettier. You might decide to weigh them. You could count the number of standalone chunks (words). And so on.
The good news is these are, in fact, relevant differences. The bad news is because you’re unlettered, there’ll always be certain differences you won’t pick up, like choice of words, context, and individual letter differences — or differences you highlight that aren’t actually differences, e.g. synonyms.
This is pretty much how humans learnt about ourselves. At first, we used symptoms to differentiate sick and healthy, then organ damage, then cell differences. While these were undeniably useful, we were barely scratching the surface. Only a few decades ago did we discover that there are literal smaller building blocks (ATCG) that exist in a given order and could differentiate these two major groups and explain many complicated observations.
Sequencing technology is how we define the order of the building blocks.

What does “cancer genomics” actually mean?
Technically and narrowly speaking, being a cancer genomics researcher means that you:
- interact with and are intimately familiar with cancer genome sequence data from one or more sequencing technologies
- apply fit-for-purpose bioinformatics tools to transform that data into a list of variants relative to a reference or ground truth genome
- label these variants with genomic elements
- interpret their implications using knowledge bases, peer-reviewed literature, and prior experience
Ideally — and again, narrowly — this is where genomics analysis stops. Genomics studies the entire genome (~3 billion letters), which is what you’ve analysed at this point. It also stops here because that’s as much information as you can mine from what is still, as of 2026, the gold standard of clinical genome sequencing: short-read sequencing (Illumina).
What short-read sequencing has done well
With short-read sequencing, the scientific and cancer research community has achieved a lot. We have:
- identified and classified variants and mutations
- defined molecular subtypes of multiple cancers
- informed clinical management and contributed to positive outcomes
But science progresses by acknowledging limitations alongside success. Some of those limitations directly motivated the rise of long-read sequencing, such as nanopore. To appreciate why, it helps to start with where short reads fall short.
Key limitations of short-read sequencing
1. Read length and structural variation
The output from short-read technology is short, with final read lengths often not exceeding 300 bp. When you compare 300 to 3 billion, the gap is disorienting.
In practical terms, this meant it was difficult to confidently detect large variants larger than 300 bp. If such variants were influencing a phenotype or symptom, we could easily miss them.
With nanopore sequencing, reads can reach several million base pairs. While still shorter than the full genome, this is orders of magnitude closer to 3 billion than 300. As a result, large insertions, deletions, rearrangements, and complex structural variants become much more tractable.
2. DNA methylation as a separate problem
On top of the ATGC sequence of the genome, there exists another layer of information: DNA methylation. This is a chemical property that does not appear as a letter change, but typically manifests on cytosines (and sometimes adenines).
With short-read sequencing, methylation required a separate, chemically destructive process and did not support whole-genome interrogation. This limited researchers to incomplete and often mediocre-quality data.
Nanopore sequencing, by contrast, detects methylation directly from the electrical signal during sequencing. This means that sequencing ATGC automatically extracts methylation information, yielding more complete data without additional experiments.
3. Variant–methylation disconnect
Because variant calling and methylation profiling were detached processes in short-read workflows, it was impossible to directly study variants and methylation together. Despite strong biological motivation, especially in cancer gene regulation, scientists could only infer links indirectly.
With long reads, both variant information and methylation are present on the same molecule. This enables read-level resolution and direct variant–methylation linkage. At least in theory.
Phase-aware and haplotype-aware analysis
A haplotype refers to a group of variants that are inherited together on the same physical DNA molecule. Phasing is the process of determining which variants sit together on the same haplotype.
This distinction matters because biological effects are often not driven by single variants in isolation, but by combinations of changes acting together on the same allele. Short-read data struggles with phasing because reads are too short to span multiple variants. Long reads, by contrast, can cover multiple variant sites and their associated methylation states in one continuous molecule.
With the software package developed during my PhD, I demonstrated the practicality of read-based variant–methylation linkage with nanopore data for:
- single-nucleotide variants
- insertions
- deletions
- breakend structural variants
The value of this approach is that instead of defining differences purely in terms of genomic positions — as is typical in short-read analysis — we can study patterns of co-occurrence across molecular layers. We can ask which changes happen together, on which haplotype, and what that implies for disease mechanisms. This contributes to the growing body of work that answers the question: what could we achieve with long-read sequencing?
So, is long-read sequencing “better”?
Nanopore technology is a little over 10 years old at the time of writing. In the genomics and clinical research timeline, it is not yet a mature adult.
Multiple studies have shown that long reads can detect and resolve large variants that are difficult for short reads, in both germline and somatic contexts. Nanopore has also shown utility in RNA sequencing and pharmacogenomics, both relevant to the clinic.
So why hasn’t it been widely adopted clinically? Several barriers still exist:
- Complex, non-automated workflows, which are problematic in clinical settings
- Non-standardised analytical methods make regulatory clearance difficult
- Rapidly evolving algorithms, particularly basecalling and error correction, complicate assay stability
- Throughput and cost considerations, although both are improving
Clinical assays require regulatory approval, and that stability is hard to achieve while the technology and analytical methods are still evolving. Get a glimpse of what I mean on this database for all long-read tools developed so far.
Conclusion
Long-read sequencing clearly excels in specific areas compared to short-read sequencing, particularly for structural variation, methylation, and haplotype-aware analysis. However, it must overcome workflow, standardisation, and regulatory barriers before becoming routine in clinical diagnostics.
For now, the two technologies are complementary rather than competitors, and understanding where each fits is central to modern cancer genomics.
메타데이터
- post_id
- 96bca2cc9d6b
- slug
- nanopore-vs-illumina-cancer-genomics-96bca2cc9d6b
- url
- https://medium.com/@gearthdexter/nanopore-vs-illumina-cancer-genomics-96bca2cc9d6b
- canonical_url
- https://medium.com/@gearthdexter/nanopore-vs-illumina-cancer-genomics-96bca2cc9d6b
- author_url
- https://medium.com/@gearthdexter
- status
- ok
- fetched_at
- 2026-06-09 15:37:30