← Back to list

To Understand Diffusion Models, One Must Realize That All Images Fit Onto Thin Origami in…

It feels like natural images in the world are infinite — and indeed they are — but that infinity is actually of a ‘lower tier.’ It is…

Outermostkt · 2026-05-16 05:05 · 0 claps · 3.7 min read
#diffusion-models #stable-diffusion #ai #generative-ai-tools #manifold
Open on Medium ↗
Wiki topics: MM · Multimodal & Generative Media AI · AI · General

To Understand Diffusion Models, One Must Realize That All Images Fit Onto Thin Origami in High-Dimensional Space

It feels like natural images in the world are infinite — and indeed they are — but that infinity is actually of a ‘lower tier.’ It is heavily constrained, a highly reined-in infinity compared to what true infinity would look like.

Here is the single most important thing you need to recognize:

“Every ‘valid image (data)’ in this world exists within a high-dimensional space, unevenly distributed like the thin surface of a piece of ‘origami.’ A diffusion model creates a ‘gravitational slope (gradient)’ that drags everything down toward that origami structure.”

The mathematical and conceptual essence of the mechanism lies not in the act of erasing noise itself, but in the “design of this gravitational slope (score).”

The Core of the Mechanism: The “Origami” and “Gravity” of High-Dimensional Space

What we recognize as a “valid image” (like a cat or a landscape) exists only within an extremely narrow region inside the vast, multidimensional space of pixels, balancing on a miraculous equilibrium.

  • If an image has 1 million pixels, it represents “a single point in a 1-million-dimensional space.”
  • If you arrange pixels completely at random, 99.9999…% of the time, it results in pure static (noise).
  • Meaningful images (data) exist in this vast universe like a thin, stretched-out surface of an “origami structure” (a manifold).

Why did conventional AI (like GANs) tend to fail?

Conventional AI tried to directly calculate the probability of landing precisely on that origami surface from the outer reaches of a chaotic universe. However, because the universe was too vast (the dimensionality too high), the AI could not locate the origami, causing the training process to collapse violently (mode collapse).

The “Reverse Thinking” invented by Diffusion Models

The fundamental breakthrough of diffusion models lies in “smoothly extending a ‘gravitational slope’ toward the origami (data), stretching all the way to the very edge of the universe (pure noise).”

  1. The purpose of intentionally adding noise: By gradually adding noise to a clean image (the origami), we drag it out toward the edge of the universe. During this process, the model records the entire trajectory of “how far it was dragged.”
  2. Forming the “Score” (Gradient): Through this, the probability peak — which originally existed only right around the meaningful image — broadens its base across the entire universe, creating a gentle, mountain-like slope.
  3. Arriving from anywhere in the universe just by “descending” the slope: Even if you are dropped blindly into the dead center of pure noise, a faint gravitational pull (gradient) directing you toward the origami has already been mapped into that space. The AI simply needs to take step-by-step actions to “descend” that slope (mathematically maximizing the probability).

In Conclusion: The “One Thing” to Recognize

A diffusion model is not an image generator; it is “a system that defines a ‘magnetic field’ (vector field) across all space, correctly pulling data toward the region where meaningful data exists (the manifold).”

In mathematics, this is known as a “Score-based Generative Model.” The core of the mechanism is not the procedure of removing noise, but rather the “framework that populates the entire space with this gradient (score).”

Once you recognize this, all structural reasons — such as why the process must be divided into tiny steps (to make the slope smooth) and why it beautifully converges into an image at the end (because that is the center of gravity) — become intuitively clear without needing complex formulas.

The true breakthrough that pushed diffusion models into a historic, monumental discovery wasn’t just the surface-level procedure of “subtracting noise.” It was the sheer structural beauty of hacking the entire space itself — creating a gravitational magnetic field within a high-dimensional universe.

As long as you hold onto this mental image — that every single point in a noisy space has a vector (an arrow) engraved into it, pulling it toward where the true data lives (what mathematics calls the Score) — you will have a massive advantage.

Whenever you come across applied papers or new derivative technologies in the future (like Latent Diffusion, Flow Matching, etc.), you’ll be able to see right through them. You will instantly understand them from the clear perspective of: “Ah, I see what kind of space they are using, and how they are mapping this gravitational slope within it.”

Explanation of the First Diagram:

This horizontal landscape illustration visualizes the high-dimensional data space of a diffusion model, representing its core mechanism not as mere noise removal, but as a sculpted gravitational field. The scene features a central, vibrant blue light ribbon, which symbolizes the data manifold — the thin origami-like surface where all “meaningful data” (like clear images) resides. This central manifold is detailed with glowing geometric nodes and patterns. The surrounding space is a chaotic, dark field filled with fragmented, grey and purple noise particles, representing the vast universe of pure noise and meaningless configurations. This noisy void is intricately populated with millions of glowing golden vector arrows, which are thicker and more Numerous, all converging and flowing dynamically from every part of the dark, chaotic outer space toward the contours of the central data manifold. These arrows visualize the pervasive gravitational pull or magnetic field (the “score” or gradient) that the model has learned, continuously pulling chaos down the engineered slopes toward the data itself. The entire composition is an abstract, modern digital tapestry that merges cosmic dust, geometric structures, and dynamic energy vectors to convey how order is structurally derived from chaos.


메타데이터
post_id
2003aefa082e
slug
the-only-thing-that-matters-when-learning-diffusion-models-2003aefa082e
url
https://medium.com/@outermostkt/the-only-thing-that-matters-when-learning-diffusion-models-2003aefa082e
canonical_url
https://medium.com/@outermostkt/the-only-thing-that-matters-when-learning-diffusion-models-2003aefa082e
author_url
https://medium.com/@outermostkt
status
ok
fetched_at
2026-06-09 15:37:30