How do we Actually Sequence DNA?
Disclosure: I wrote this article and used AI as a “professional editor” to refine and shorten overly long or complicated sentences.
How do we Actually Sequence DNA?

Image generated with the assistance of AI.
Disclosure: I wrote this article and used AI as a “professional editor” to refine and shorten overly long or complicated sentences.
This article is less about a direct clinical application and more about a technology that made much of the recent progress in biotechnology and biomedicine possible. The shift in healthcare toward more personalized medicine depends on many different innovations and achievements, such as CRISPR for gene therapy or the successful delivery of messenger RNA in mRNA vaccines. But DNA sequencing provides the foundation that ties many of these advances together.
Advances in sequencing technology have been fundamental for the development of personalized medicine. At the same time, the history of sequencing provides a very intuitive demonstration of just how quickly biotechnology has progressed and how far we have already come.
The National Human Genome Research Institute defines DNA sequencing as “a general laboratory technique for determining the exact sequence of nucleotides, or bases, in a DNA molecule.” This sounds simple and straightforward. But how do we actually determine a DNA sequence if we cannot directly see the nucleotides? They cannot be distinguished directly, not with conventional microscopy, and even the highest-resolution imaging methods cannot reliably identify individual bases in a way that enables sequencing.
This article will look at the early beginnings of DNA sequencing and compare them with modern sequencing technologies, highlighting the remarkable advances made over the past decades.
As with the other articles in this series, the explanation is intentionally simplified. The only background knowledge needed is that our genome is made of DNA (deoxyribonucleic acid), which consists of four building blocks called nucleotides: adenine (A), thymine (T), cytosine ©, and guanine (G).
DNA is double-stranded, and the two strands are complementary to each other. This means that if we know the sequence of one strand, we automatically know the sequence of the other. Adenine always pairs with thymine, and cytosine always pairs with guanine, meaning that if one strand has cytosine, the other strand has guanine at that position.
The “godfather” of sequencing: Sanger Sequencing
The earliest DNA sequencing methods were developed in the 1970s, and interestingly two different approaches emerged around the same time. One method was developed by Allan Maxam and Walter Gilbert, and the other by Frederick Sanger. While both methods were important early breakthroughs, the approach developed by Sanger quickly proved more practical and ultimately dominated DNA sequencing for more than 30 years. Even today, Sanger sequencing is still widely used.
The key idea behind Sanger sequencing is to copy a DNA strand while occasionally inserting special nucleotides that stop the copying process. By examining where these stops occur, the DNA sequence can be reconstructed. But how does this work in practice?
The process begins with a primer, a short piece of DNA (usually about 20–30 nucleotides long) that binds to a specific location on the DNA strand to be sequenced through complementary base pairing (adenine binds thymine, and cytosine binds guanine).
Because of this complementarity, primers can be designed to bind very precisely. Even within the human genome, which contains more than 3 billion base pairs, a primer can be designed that binds to only one specific location.
Because DNA is double-stranded, the strands first need to be separated. This is done by heating the DNA to about 95°C (203°F). At this temperature, the bonds between the two strands break, and the DNA becomes single-stranded.
Next, the temperature is lowered to around 50–60°C, which allows the primer to bind to its target sequence on the template strand.
Once the primer is bound, an enzyme called DNA polymerase begins extending the primer by adding nucleotides that match the template DNA strand. In this way, a new DNA strand is created that begins at the primer and copies the sequence of the original DNA strand. The new sequence is complementary to the original template. This step usually occurs at around 70°C (158°F).
The key innovation of Sanger sequencing is the addition of special nucleotides that terminate DNA synthesis. These modified nucleotides lack the chemical structure required to attach another nucleotide. When one of these nucleotides is incorporated, DNA synthesis stops.
Because termination occurs randomly, the reaction generates fragments that collectively represent every possible stopping point along the DNA sequence. As a result, the reaction produces a collection of DNA fragments that all start at the same primer but end at different points along the DNA sequence.
Originally, Sanger sequencing required four separate reactions, each containing one type of terminating nucleotide. This ensured that every fragment produced in that reaction ended with the same base.
Once the reactions are complete, the resulting DNA fragments are separated by size using polyacrylamide gel electrophoresis. The gel acts like a very tight mesh. Short DNA fragments can move through this mesh more easily and therefore travel faster and farther, while longer fragments encounter more resistance and move more slowly.
The DNA fragments are placed at the top of the gel, and an electric field is applied. Because DNA carries a negative charge, the fragments move through the gel toward the positive electrode at the bottom.
Since all fragments start at the same position but end at different points along the DNA sequence, they form a pattern of bands separated by length. By comparing the band patterns from the four reactions, the DNA sequence can be read one base at a time.
The shortest fragments appear closest to the bottom of the gel and represent the first bases in the sequence. Progressively longer fragments appear higher in the gel and correspond to later positions in the DNA strand.
Sanger sequencing converts the invisible sequence of DNA into a visible ladder of fragments, where each band represents a specific nucleotide position that can be read from bottom to top. To read the image below, focus on the four lanes, each labeled with one DNA letter (A, T, C, G). Each horizontal band marks where a DNA fragment stopped growing. Start at the bottom, the shortest fragments are there, and read upward. The lowest band across all lanes is in the A lane, so the first base is A. The next band up is in the T lane, so the second base is T, and so on. Reading bottom to top gives the full sequence: A — T — C — G — A — T — G — C -T — A.

Image generated with the assistance of AI.
The Human Genome Project Used Sanger Sequencing
Performing all the steps described above is time- and labor-intensive. Early Sanger sequencing required four separate reactions for each DNA sequence and careful preparation of polyacrylamide gels. Each gel could accommodate roughly 8–16 samples (since each sample required four lanes) and typically had to run for several hours, often overnight. This workflow was accurate but slow and clearly not designed for large-scale sequencing.
A major improvement came with the development of capillary Sanger sequencing. While the fundamental concept remained the same, two key innovations made the method far more automated and scalable.
First, the four chain-terminating nucleotides were labeled with different fluorescent dyes, allowing all four reactions to be combined into a single sequencing reaction rather than four separate ones.
Second, DNA fragments were separated inside thin capillary tubes filled with a polymer matrix that acts like a microscopic sieve, replacing the traditional polyacrylamide gel.
The principle of separation remained the same: shorter DNA fragments move faster through the polymer matrix than longer ones. As the fragments pass a detection window inside the capillary, a laser excites the fluorescent dye attached to the terminating nucleotide. The color of the emitted light reveals which nucleotide ended the fragment, and a computer reconstructs the DNA sequence automatically.
The result is a colorful wave pattern called a chromatogram. Each colored peak represents one nucleotide, and the sequence is simply read left to right.

Image generated with the assistance of AI.
Instruments were soon developed that could run dozens of capillaries in parallel, greatly increasing sequencing throughput while reducing manual work. This technological progress was essential for the Human Genome Project, launched in 1990, which relied largely on capillary Sanger sequencing.
Large sequencing centers operated rooms filled with automated sequencers running continuously. Thousands of scientists worked for more than a decade to complete the project. The first draft of the human genome was published in 2001, with a more complete reference sequence published in 2003.
The project required sequencing and assembling roughly three billion DNA base pairs. The Human Genome Project required an estimated $3 billion USD and more than a decade of work to complete a single human genome. Today, the same task can be accomplished in a matter of days for roughly $500–$1,000 using modern sequencing technologies.
So what changed?
The Game Changer: Next-Generation Sequencing (NGS)
The key limitation of Sanger sequencing is that it reads one DNA fragment at a time. Even highly automated capillary systems still process individual sequencing reactions. This makes the method highly accurate, but difficult to scale.
Next-generation sequencing (NGS) changed this completely by shifting from sequential to massively parallel sequencing. Instead of reading one DNA fragment at a time, modern sequencing systems read millions to billions of fragments simultaneously.
While several NGS technologies were developed, the most widely used approach today is sequencing by synthesis (SBS). At its core, the principle remains similar to Sanger sequencing: DNA polymerase builds a new DNA strand by incorporating nucleotides that are complementary to a template. The key difference is how this process is scaled and detected.
One key difference is that the DNA needs to be prepared before it can be sequenced. The DNA fragments that are sequenced are typically 200–600 base pairs long. Longer DNA, such as an entire genome, must first be fragmented. A short piece of DNA, called an adapter, is then added to each fragment, which later enables the fragment to be captured and positioned on the reaction surface.
The prepared DNA fragments are then loaded onto a specialized surface called a flow cell. The surface of the flow cell contains billions of tiny reaction sites, each containing short DNA sequences (oligonucleotides) that are complementary to the adapter sequences. These oligonucleotides capture the DNA fragments and anchor them at defined locations on the surface.
Once attached, each fragment is locally amplified to form a small cluster of identical DNA copies. This amplification step ensures that the signal generated during sequencing is strong enough to be detected reliably.
Sequencing then proceeds in cycles. During each cycle, nucleotides are added to the growing DNA strand. Similar to the chain-terminating nucleotides used in Sanger sequencing, these nucleotides temporarily block further extension after incorporation, meaning that only one nucleotide can be added in each cycle. Unlike Sanger sequencing, however, this modification is reversible.
Each incorporated nucleotide also carries a fluorescent label. After a nucleotide is added, the flow cell is imaged, and the color of the signal reveals which base was incorporated at each position. The fluorescent label and blocking group are then removed, allowing the next cycle of sequencing to proceed.
By repeating this process cycle after cycle, the DNA sequence of every fragment can be determined one base at a time. Because all fragments are sequenced in parallel across the flow cell, an enormous number of sequences are generated simultaneously.
Rather than producing a single continuous read, NGS generates millions to billions of short DNA sequences, known as reads. To reconstruct the original DNA sequence, computational methods are used to identify overlaps between these reads and assemble them into a complete sequence, much like solving a complex puzzle.
Next-generation sequencing (NGS) transformed genomics. What once required years of work and enormous sequencing facilities during the Human Genome Project can now be accomplished in a matter of days using a single instrument at a fraction of the cost.

Third-Generation Sequencing
While next-generation sequencing dramatically increased throughput by reading many short DNA fragments in parallel, third-generation sequencing takes a different approach by reading much longer stretches of DNA, often tens of thousands of bases or more in a single read. This is commonly referred to as long-read sequencing.
Two main technological approaches have emerged. One method observes DNA polymerase as it copies a DNA strand in real time, detecting each nucleotide as it is incorporated. Another method passes a DNA molecule through a tiny pore and measures changes in electrical current to determine the DNA sequence. In both cases, sequencing occurs in real time, without the repeated imaging cycles required in short-read sequencing.
The ability to generate long reads offers important advantages. Long-read sequencing can span complex or repetitive regions of the genome that are difficult to reconstruct from short fragments. It also enables more accurate detection of structural variations, such as insertions, deletions, and rearrangements, which play a key role in many diseases.
Final Thought
While sequencing technologies are often described in generations, all three approaches — Sanger sequencing, next-generation sequencing, and third-generation sequencing — are still actively used today. Each method has strengths that make it best suited for specific applications. Sanger sequencing remains the gold standard for small-scale, highly accurate validation. Next-generation sequencing enables large-scale studies by reading millions to billions of fragments in parallel. Third-generation sequencing provides long reads that help resolve complex regions of the genome that are difficult to reconstruct from shorter fragments. Together, these technologies form a complementary toolkit that continues to drive advances in biology and medicine.
If you’d like to keep reading:
Subscribe to my newsletter on **Substack** (free, email delivery)
If you’d like to follow my work professionally (and read articles there as well):
Connect with me on **LinkedIn**
If you find this writing useful and want to support independent science communication:
You can support my work on **Patreon** (optional)
메타데이터
- post_id
- 322653ebe894
- slug
- how-do-we-actually-sequence-dna-322653ebe894
- url
- https://medium.com/@scienceworthknowing/how-do-we-actually-sequence-dna-322653ebe894
- canonical_url
- https://medium.com/@scienceworthknowing/how-do-we-actually-sequence-dna-322653ebe894
- author_url
- https://medium.com/@scienceworthknowing
- status
- ok
- fetched_at
- 2026-06-09 15:37:30