← Back to list

[4D DNA Decoding] The 98% Problem: Why the Human Genome Is Not Mostly Junk

# The 98% Problem: Why the Human Genome Is Not Mostly Junk

이영재 · 2026-05-20 00:53 · 0 claps · 5.2 min read
#4d-dna-decoding #junk-dna #human-genome #genome #decoding-dna
Open on Medium ↗
Wiki topics: GEN · Genomics & Sequencing GNM · Genome · General 💻 · Programming

[4D DNA Decoding] The 98% Problem: Why the Human Genome Is Not Mostly Junk

The 98% Problem: Why the Human Genome Is Not Mostly Junk

Article 1 of 16 in the series “The 4D Blueprint Inside the Genome: A Jamming-Lattice Reading of DNA.”

TL;DR. The label “junk DNA,” coined by Susumu Ohno in 1972 for the ~98% of the human genome that does not code for proteins, was always an admission of ignorance, not a finding. In 2012 the ENCODE consortium documented biochemical activity across roughly 80% of the genome. This series argues for a stronger claim: the non-coding fraction is a 4D placement specification — the structural blueprint that decides how the coding 1.5% is read, where, when, and in what 3D context. Once that placement is locked, growth produces a unique solution.

— -

What is “junk DNA”?

Junk DNA is the colloquial name for non-coding DNA — the parts of a genome that do not directly specify protein amino-acid sequences. The term was popularized by geneticist Susumu Ohno in his 1972 Brookhaven Symposium paper ”So much ‘junk’ DNA in our genome,” where he argued that mammalian genomes are largely (~90%) non-functional accumulations of past evolutionary debris.

The label stuck for two reasons. First, only a tiny fraction of the genome codes for proteins, and protein-coding was the most legible function biology had at the time. Second, large portions of the non-coding genome are made of repeats (transposable elements, simple repeats, satellite DNA) that look like duplication on a massive scale.

Both observations are correct on the surface and misleading underneath.

How much of the human genome is actually protein-coding?

About 1.5%. The human genome is roughly 3.1 gigabases. Protein-coding exons account for ~1.5% of that. The remaining ~98.5% is non-coding: introns (~26%), pseudogenes (~1–2%), transposable elements (~45%), regulatory sequences (promoters, enhancers, insulators), non-coding RNA genes, and repetitive structural elements.

A useful contrast: the C-value paradox notes that genome size does not track organismal complexity. The human genome is ~3.1 Gb. The onion genome is ~16 Gb. The lungfish genome exceeds 130 Gb. If genome size were proportional to “what we are,” none of this would make sense. The 1.5% coding fraction is stable across complex eukaryotes; everything else varies dramatically — exactly the pattern we should expect if the non-coding fraction is doing something structural and contextual rather than coding.

What did the ENCODE Project actually claim in 2012?

In September 2012 the ENCODE Consortium published 30 coordinated papers in Nature, Genome Research, and Genome Biology. Their headline number: ~80.4% of the human genome has at least one detectable biochemical activity — transcription, transcription-factor binding, chromatin signature, DNase hypersensitivity, histone modification, or DNA methylation.

ENCODE catalogued, among other things, more than 70,000 promoter-like regions and more than 400,000 enhancer-like regions. These are not random sites. They are organized, they are reproducible across cell types, and they explain why the same DNA produces a liver cell here and a neuron there.

ENCODE’s 80% figure has been contested. Critics correctly point out that “biochemical activity” is not the same as “selected function,” and that some activity may be background noise of an active nucleus. That methodological debate is real and continues. But it changed the burden of proof. After 2012 the default assumption became “this region does something until shown otherwise,” not “this region is junk until shown otherwise.”

The reframe: from noise to placement specification

Here is the thesis of this series, stated once at the start and earned over the next fifteen articles.

The ~98% non-coding fraction of the genome is a 4D placement specification. It is not redundant noise. It is the structural blueprint that decides where coding sequences are pinned in 3D space, in which mechanical regime they sit, which regulatory partners they are wired to, and in what temporal order their activity gets locked in.

The proper analogy is not “code with junk around it.” The proper analogy is a printed circuit board. The traces, vias, ground planes, mounting holes, mechanical supports, and thermal pads occupy most of the board area. They are not the chips — but without them, the chips do nothing useful. They specify where and how the chips connect, dissipate heat, and resist mechanical stress.

A genome is the same kind of object. The coding 1.5% is the chips. The non-coding 98.5% is the board.

Why a 1D string of letters cannot, by itself, build an organism

Three observations make this concrete.

1. The same DNA produces different cells. Every cell in your body has effectively identical DNA, yet a liver cell, a neuron, and a skin cell are profoundly different. The difference is not in the sequence. It is in which regions are accessible, which 3D loops are formed, which regulatory elements are bound, and which mechanical compartment the chromatin sits in. That information is non-coding.

2. Genome size and organism complexity decouple. The C-value paradox is real. If 1D sequence content were the unit of biological information, genome size should track complexity. It does not.

3. Deletion of “junk” regions causes real defects. Over the last decade, removal of large non-coding regions — including repeat-derived enhancers, long non-coding RNA loci, and CTCF binding sites at TAD boundaries — has repeatedly produced developmental and disease phenotypes. The “junk” is load-bearing.

Each is independently sufficient to retire “junk DNA” as a literal description. Combined, they force a stronger conclusion: the 98% is doing structural work that no 1D reading can capture.

So what is the 4D blueprint, concretely?

In this series I will use the framework developed in a deterministic-interpretation whitepaper for the mouse genome mm39 (GRCm39 assembly, Zenodo DOI 10.5281/zenodo.17963127). The framework names four primitives:

  • Shells — contiguous regions of similar mechanical stiffness, computed from GC content, CpG density, and AT-tract density.
  • Anchors — positions where stiffness changes abruptly. They live at region edges, shell boundaries, and shell cores.
  • Motors — transcription start sites, where activity is injected into the structure.
  • Loops — couplings between motors and nearby anchors, with explicit base-pair distances.

Together these form an object called A4 — a 4D arrangement that any region of any genome can be deterministically reduced to. In the next five articles I will walk through what each primitive is, how it is computed, and why these four are sufficient.

Above A4 sits a physical layer: the rigid shell (a jammed lattice in the source physics whitepaper, DOI 10.5281/zenodo.17932567). Once A4 is locked into a jammed configuration, the developmental trajectory is forced toward a unique solution. This is the strong claim — that placement does not merely influence outcome, it selects it.

What this series will and will not claim

This series will argue, with citations to three DOI-registered whitepapers, that:

  1. The non-coding ~98% is a structural specification, not noise.
  2. A small fixed vocabulary (shells, anchors, motors, loops) is sufficient to write that specification.
  3. A small fixed grammar (four operations: INIT, SCONSERV, SDISSIP, JEVENT) is sufficient to describe how that specification unfolds in time.
  4. The framework is deterministic, auditable, and reproducible under a strict LOCK → Derive → Gate contract.

This series will not predict biological function in arbitrary detail, will not offer medical advice, and will treat dating estimates and untested external theories as outside scope. Every claim links back to an artifact you can download, verify by sha256, and rerun.

— -

Sources & verification

Author: Young Jae Lee (ORCID 0009–0002–7535–8245). No funding, no conflict of interest. This article reports only LOCK-derived claims from the cited whitepapers; no clinical advice.

Next in the series: From 1D Letters to 4D Placement — The Four Primitives. We install the lens: shells, anchors, motors, loops, and why exactly these four.


메타데이터
post_id
8ce7fa231231
slug
4d-dna-decoding-the-98-problem-why-the-human-genome-is-not-mostly-junk-8ce7fa231231
url
https://medium.com/@rego093/4d-dna-decoding-the-98-problem-why-the-human-genome-is-not-mostly-junk-8ce7fa231231
canonical_url
https://medium.com/@rego093/4d-dna-decoding-the-98-problem-why-the-human-genome-is-not-mostly-junk-8ce7fa231231
author_url
https://medium.com/@rego093
status
ok
fetched_at
2026-06-09 15:37:30