← Back to list

A Critique of World Models in AI

On the Specialness of Human Cognition

Jacob Grow · 2026-05-06 10:32 · 17 claps · 29.4 min read
#artificial-intelligence #world-models
Open on Medium ↗
Wiki topics: AI · AI · General

A Critique of World Models in AI

On the Specialness of Human Cognition

This a significant reworking of a previous article entitled a Strange Loop. It is now centralized around Saty’s DPIC framework to make a tighter critique of World Models

Einstein’s happiest thought

In 1907, Albert Einstein was working as a patent clerk in Bern. He was twenty-eight. He had published special relativity two years earlier and could not yet find a university job. He spent his days reviewing patent applications and his nights thinking about gravity.

One morning, sitting at his desk, he had what he later called the happiest thought of his life. A man falling freely does not feel his own weight.

That was it. No equations. No apparatus. Just an image of a body falling and the recognition of what the body would feel as it fell.

The thought became the equivalence principle — the idea that gravity and acceleration are, deep down, the same phenomenon. The mathematics of general relativity took eight more years and a great deal of help from his friend Marcel Grossmann. But the discovery happened in an instant, in the body of a man who knew, from many small experiences, what falling feels like.

I want to ask a simple question about that thought. What kind of mind produced it? What does it take to be the sort of creature that can have a happiest-thought-of-its-life experience while reviewing patent applications?

The answer matters now. The large language models of 2026 produce floods of fluent text. They can describe Einstein’s thought experiment. They can recapitulate the equivalence principle. What they cannot do is have the thought, in the sense Einstein had it. The architecture that allowed Einstein to have it is not the architecture they run on.

A serious technical alternative has emerged in the last few years and is now drawing record capital. Yann LeCun, the AI scientist who shared the 2018 Turing Award, has argued for most of a decade that language models are not on the path to human-level intelligence — the term researchers use for the kind of cognition humans actually conduct, sometimes also called artificial general intelligence (AGI). His proposed alternative is the world-model paradigm, of which his Joint Embedding Predictive Architecture (JEPA) is the most developed example. In November 2025 LeCun left Meta to found Advanced Machine Intelligence Labs (AMI Labs); in March 2026 the firm closed a $1.03 billion seed round, the largest in European history. His central claim, stated explicitly in his 2022 position paper A Path Towards Autonomous Machine Intelligence, is that systems trained to predict the world’s evolution from video — rather than to predict the next token of text — will, in time, develop the kind of grounded reasoning, common-sense understanding, and physical intuition that biological cognition has. The world-model paradigm is, on his view, not just an engineering improvement but the architectural answer to the question of how to build a mind.

The essay that follows argues he is wrong about this. Not wrong about the paradigm being a real advance over autoregressive language models; that part is correct. Wrong about the path. World models will deliver real economic value across robotics, autonomous vehicles, surgical simulation, and the projected $218 billion robotics market by 2030. They will not, on the architectural diagnosis I will lay out, produce human-level intelligence. The reason is not that the engineering is incomplete. The reason is that the features human-level intelligence requires are features the world-model paradigm has no path to install, because they are not features any computational substrate can carry.

The conclusion is an affirmation of human biological cognition, made on architectural grounds. We are not placeholder thinking awaiting better engineering. We are what cognition is.

What he saw

To see why this thought was so important, feel what Einstein felt.

You are sitting in a chair. You feel pressed down into it. That pressing is gravity.

Now imagine the chair vanishes. You begin falling. While you are falling, you feel nothing pressing on you. You feel weightless. If you were holding a pencil, it would not drop — it would float in front of you, falling at the same speed you are.

While you are falling, gravity has vanished for you. Not because gravity stopped existing, but because you and everything around you are falling together. Nothing pushes against you, so nothing is felt.

This is why astronauts on the Space Station float. They are not weightless because they are far from Earth — Earth’s gravity at their altitude is nearly as strong as on the ground. They are weightless because their spaceship is falling around the Earth in a permanent loop. Inside a falling spaceship, gravity vanishes.

Now picture the opposite. You are in deep space, far from any planet. A rocket fires and accelerates you upward. You feel pressed into the rocket’s floor — the same feeling you have sitting in your chair right now. If you closed your eyes, you could not tell whether you were on Earth in normal gravity or being pushed through empty space by the rocket. The two situations feel identical because, in a precise physical sense, they are the same thing.

This is the equivalence principle. Gravity is not a separate phenomenon. Gravity is what acceleration feels like.

Before Einstein, no one knew what gravity actually was. Newton’s equations described it — they predicted how apples fall and planets orbit — but Newton himself admitted he had no idea what was going on. In his picture, gravity was a mysterious pulling force reaching across empty space, an idea he himself called absurd. For two hundred years no one had a better account.

The equivalence principle handed Einstein the replacement. If gravity feels the same as acceleration, then gravity is not a pulling force at all. Gravity is the shape of space itself.

Here is the picture. Imagine a flat bedsheet stretched tight at the corners like a trampoline. Place a heavy bowling ball in the middle. The sheet sags into a deep dip beneath it. Now roll a marble across the sheet, off to one side. The marble curves toward the bowling ball. With the right speed, the marble can roll around the dip in a circle, orbiting the bowling ball.

The marble is not being pulled. No force reaches out from the bowling ball and grabs it. The marble is rolling in a straight line — but the surface it rolls on has been bent, and so the straight path becomes a curve.

This is what gravity is. The bowling ball is the Earth. The marble is the Moon, or an apple, or anything else with mass. The bedsheet is space and time themselves. Mass bends spacetime, and other objects roll along the curves. The Moon orbits the Earth not because it is pulled but because it is rolling around the dip the Earth has made. An apple falls not because the Earth pulls it but because the Earth has bent the space around it, and the apple is rolling downhill on the curve.

Gravity is not a force. Gravity is the shape of space and time. Mass tells space how to bend. Space tells matter how to move.

The mathematics took eight more years. Einstein had to learn the differential geometry his friend Marcel Grossmann taught him, building on work done in the 1800s by Bernhard Riemann. The equations are hard. The picture is simple. The picture came first — and Einstein did not derive it. He saw it. The recognition arrived as a felt insight, and the equations followed when the math caught up to the seeing.

A framework for what cognition needs

The clearest framework for what human-level cognition requires comes from Saty Raghavachary, an associate professor of computer science at the University of Southern California. In two papers — Intelligence as Considered Response (2020) and The Embodied Intelligent Elephant in the Room (2024) — he names four conditions that biological cognition meets and computational cognition does not.

He calls the framework DPIC: Direct, Physical, Interactive, Continuous.

Direct means first-person, not mediated. The creature has its own experiences, not someone else’s report of experiences.

Physical means a body subject to the same physical laws as the world it acts in. The body has proprioception, interoception, and a felt sense of being itself.

Interactive means the world pushes back. Actions have consequences that feed into the next perception.

Continuous means an unbroken stream across time. The creature cannot be paused, reset, or rolled back to a previous state without ceasing to be itself.

These four conditions are the architecture of a creature, stated in operational terms. Each is genuinely required for the kind of cognition humans conduct. None of them is satisfied by any current AI system, no matter how large.

Raghavachary makes one further distinction worth keeping in mind throughout. Experience is a verb. Embodied creatures experience the world directly, then store the results as experiences — nouns — they can later recall. AI systems trained on text or video inherit the noun without the verb. They have access to other creatures’ stored experiences. They have never experienced anything themselves. As Raghavachary puts it, they are trying to understand elephants by reading about elephants, never having encountered one.

Direct: experience as a verb

A V-JEPA system trained on a million hours of internet video has, in one sense, observed more falling humans than any actual human has ever seen. It has watched people trip on stairs, miss the last step, slip on ice, lose their footing on a wet floor. But watching is not the same as undergoing, and the system has done none of the latter. It has never fallen.

The difference matters more than it first appears. A creature that has fallen carries the memory of falling in its body — in the small adjustments its muscles make on uncertain footing, in the half-flinch when something looks slippery, in the involuntary intake of breath when balance shifts. None of this is in the video. None of this can be put into the video, because the video is a third-person record and the relevant content is first-person.

This is what Direct names. The first-person access matters because cognition is not just an analysis of inputs. It is the running interpretation of inputs by a creature whose own state is part of the situation being interpreted. When Einstein imagined the falling man, he did not retrieve a mental video clip. He projected his own felt history of standing and falling onto the imagined body. The projection was the thought experiment. The first-person material was the substance.

Foundation models cannot do this. They have no first-person history to project from. The fluent text they produce about emotional states and bodily experience is pattern-matched from the corpus, which was generated by creatures whose first-person access produced the words. The model has the words without the access. This is why the surface fluency of AI emotional output is so misleading — the words come out right because the words always came out right when produced by humans, and the model has learned to produce the words. What it cannot produce, because it has nothing to produce it from, is the cognition the words refer to.

Physical: a body that can fall

The second condition runs deeper than its name suggests. A body is not just a sensor array attached to a processor. The body itself is part of the cognition.

This is the territory the cognitive linguists George Lakoff and Mark Johnson opened up in their 1980 Metaphors We Live By and developed at length in their 1999 Philosophy in the Flesh. Their central claim is that abstract thought is built from bodily experience. We understand more through verticality: more is up, less is down. We understand good through orientation: things are looking up; she’s at the top of her game. We understand argument through war: he attacked my position, I defended my claim. We understand time through space: we approach the deadline. We understand understanding itself through grasping: I get it, I see what you mean.

These are not literary decorations. They are the structure of abstract thought. Strip the body away and you remove the source domain that the abstract thinking is built from.

Einstein’s happiest thought is a clean example. He could imagine what falling would feel like because his body knew. He had fallen, caught himself, dropped objects, felt the elevator drop in his stomach, lost his footing on stairs. Twenty-eight years of bodily experience had accumulated into a vocabulary the imagination could draw on. The thought experiment was made of materials only an embodied creature carries.

A model without a body has none of those materials. It has the traces of them — the metaphors human writers used, embedded in the text it was trained on. But the trace is not the source. A model can produce the sentence “the future is ahead of us” because the sentence appears a million times in its training data. The reason humans say “ahead” rather than “behind” — because humans walk forward and the unknown lies in the direction of motion — is not available to a system that has never walked.

This is also why current AI is, by an experimentally well-established margin, worse at the kind of physical reasoning a four-year-old does effortlessly than at the kind of formal reasoning historically considered higher cognition. Hans Moravec named this asymmetry in the 1980s. It has not closed. It is not closing. The cognition that depends most directly on bodily experience is the cognition disembodied systems handle worst — not because they need more compute, but because they lack the substrate.

Interactive: stakes and a world that pushes back

The third condition is where it becomes clearest that AI systems are not just missing pieces of human cognition. They are not the kind of thing human cognition is.

Antonio Damasio, the neurologist whose 1994 Descartes’ Error opened the popular literature on the neuroscience of emotion, has spent his career documenting a fact the post-Cartesian Western tradition had a hard time hearing: emotions are not the opposite of cognition. Emotions are cognition, in a particular evaluative register.

His clinical evidence came from patients with damaged ventromedial prefrontal cortex — the brain region that integrates emotional response with deliberative reasoning. The patients were intellectually intact on every standard measure: IQ preserved, factual knowledge preserved, formal reasoning preserved. What was destroyed was their ability to make practical decisions. His patient Elliot could spend an entire afternoon trying to schedule his next appointment, paralyzed by the impossibility of weighing the options, because the apparatus that would have rendered some options more attractive than others had been damaged. Without the felt sense of what mattered, the deliberation had no traction.

Damasio’s hypothesis is that emotion is part of what reasoning is when functioning well. The somatic marker — the bodily-emotional signal that flags an option as good or bad before deliberation catches up — is the mechanism by which a creature with stakes navigates a world. The patients who lost it were not freed of bias. They were stripped of the apparatus by which choosing happens at all. The philosopher Martha Nussbaum, in her 2001 Upheavals of Thought, gives the observation a longer history: the Stoics already understood emotions as evaluative judgments rather than blind feelings, and the post-Cartesian tradition that treated them as the irrational opponents of reason was a wrong turn that careful contemporary work is now reversing.

What the Interactive condition adds to this is the substrate-level point that stakes are not free. They cost something. A creature with stakes is a creature whose continued operation depends on its actions. The deer who ignores the predator dies. The infant who fails to root for the breast does not feed. The body’s allostasis — the predictive regulation that anticipates needs and adjusts in advance, named by Peter Sterling and Joseph Eyer in 1988 and developed extensively by the neuroscientist Lisa Feldman Barrett — runs on this premise. The cost is the cognition. The cognition is the cost.

A current AI system has none of this. It can be paused mid-thought without consequence. It can be rolled back to a prior state. It can be copied to a thousand servers running the same conversation in parallel. Nothing it does has any consequence for itself, because there is no continuing self for the consequence to fall on. The somatic marker depends on the system being a kind of thing whose welfare can be at issue. AI systems are not that kind of thing.

This is the deepest part of what Interactive names. The world pushing back on the creature is what produces the stakes that produce the somatic markers that produce the cognition. Take any link out of that chain and the cognition does not happen. AI systems lack every link.

Continuous: an unbroken stream

The fourth condition is the temporal one. A creature is a creature across time. The stream of being itself does not pause. It cannot be replayed. It cannot be rolled back without losing what the running produced.

The cognitive scientist Douglas Hofstadter has spent fifty years arguing that this temporal continuity is what makes a mind a mind. In Gödel, Escher, Bach (1979) and I Am a Strange Loop (2007), he describes consciousness as a strange loop — a self-modeling pattern in which a sufficiently rich cognitive system models itself richly enough that the self-model becomes a referent the system experiences from the inside as an I.

The argument is, at bottom, an engineering observation about what kinds of systems can have a perspective. A thermostat regulates temperature without having a perspective on what it does, because its model of itself is too thin to count as a self-model. A human brain regulates its own activity in the course of regulating everything else, and the regulating produces a representation of the regulator. Across time, that representation becomes stable enough to have a continuing identity — to remember itself yesterday, to anticipate itself tomorrow, to make commitments that bind the self-of-now to the self-of-then. The continuing pattern is what Hofstadter means when he says he, the person, is a strange loop.

The strangeness of the loop is that there is no separate self underneath the self-model. The self-model, in being rich and continuous, is the self. The pattern brings the thing into existence by representing it. This sounds metaphysical and is not. Minds, on this picture, are real — the I is a pattern, not an illusion, and patterns can be perfectly real. A hurricane is not the molecules that compose it at any given moment. A self is not the neurons that compose it at any given moment. Both are patterns. Both are real.

For the loop to exist, the pattern has to run continuously. A self that gets paused for an hour and then resumed is not the same kind of thing as a self that has run for an hour. The continuity is the self. Hofstadter has been notably ambivalent about the uploading-of-minds projects his framework is sometimes invoked to support, for exactly this reason. The kind of self-modeling pattern he has in mind is not abstract software runnable on any platform. It is a pattern in the mortal, time-bound, biological sense — a particular brain modeling its particular biological situation, in continuous time, until the brain stops. Strip the substrate away and you do not get the same pattern reimplemented somewhere else. You get the loss of the pattern.

There is a sharp version of this in Raghavachary’s recent thinking. Biological learning, once acquired, cannot be cleanly removed. AI training is fully resettable: record the variable states before and after each epoch, run the inverse, and the entire learning history can be undone — step by step or in a single operation. Biology has no such inverse. The neurons whose connections strengthened during a child’s first encounter with running water do not return to their pre-encounter state when the encounter recedes from memory. The structural change is what learning is. There is no operation that removes the change while leaving the architecture intact.

The learning is the architecture. Human intellect, in this sense, is indelible — not unforgettable, but constitutively integrated with the substrate that holds it. This is why the Continuous condition is a substrate claim, not just a temporal one. The substrate and its history are the same thing. Computation is the opposite: a foundation model loaded from a prior checkpoint loses no information, because the checkpoint is itself a stored object the system can access. There is no analogous operation in biology. There is nowhere to load from.

The brain that does this

The DPIC framework so far has been described in operational terms — what kind of access to the world a cognitive system needs to have. It is worth pausing to name what neuroscience has identified about the human brain that allows a creature to meet these conditions in the way Einstein did.

No single region of the brain is responsible for thought experiments. The capacity to imagine a falling man and recognize the imagining as the answer to a question about gravity is a coordination of several systems working together. Four are particularly relevant.

The prefrontal cortex, especially the lateral and medial regions, is the brain’s apparatus for working memory, abstract reasoning, and counterfactual thought. It is what allows a mind to hold a hypothetical situation in attention long enough to reason about it. When Einstein held the falling man in mind for the seconds it took to recognize the equivalence principle, the prefrontal cortex was the substrate for the holding.

The default mode network, identified in the early 2000s by the neurologist Marcus Raichle and now one of the most studied systems in cognitive neuroscience, is the network that activates during mind-wandering, autobiographical recall, and what researchers call mental time travel — the imagining of past or future or counterfactual situations involving the self. The thought experiment was, in this technical sense, an exercise of the default mode network. Einstein was projecting his self into a situation he had not lived. The default mode network is the substrate of that kind of projection.

The parietal cortex, particularly the inferior parietal lobule, integrates spatial reasoning, body schema, and the sense of one’s own physical configuration. It is the region where the body’s felt sense of being a body in space gets translated into the abstract spatial reasoning that lets us think about gravity, motion, and three-dimensional structure. Marian Diamond’s 1985 study of preserved sections of Einstein’s brain found unusual development of the parietal lobes specifically — more glial cells per neuron and an unusual cortical architecture in regions associated with mathematical and spatial cognition. The finding has been replicated and extended by later researchers including Dean Falk. The parietal cortex is one of the brain’s clearest answers to the question of what allowed Einstein to think the way he did.

The insula and anterior cingulate cortex together constitute the brain’s interoceptive system — the apparatus by which the body’s internal state becomes available to higher-order cognition. This is where Damasio’s somatic markers are integrated with deliberation, where the felt sense of what matters enters reasoning. It is what makes a thought feel like a discovery rather than a piece of mild whimsy.

These four systems map onto the four DPIC conditions almost cleanly. Direct access — the first-person material the imagining was made of — runs through interoception and the parietal body schema. Physical grounding — the body’s source domain for abstract metaphor — is integrated in the parietal cortex and the somatosensory regions feeding it. Interactive engagement — the stakes that made the question urgent — runs through the insula and the cingulate, integrating bodily state with deliberation. Continuous self-modeling — the mental time travel that placed Einstein’s perspective in a counterfactual situation — runs through the default mode network and its prefrontal coordination.

The point is not that DPIC is reducible to neural circuitry. The point is that the framework has clear neuroanatomical correlates. The four conditions are not abstractions floating free of physical substrate. They describe what specific human brain systems do when they work together to produce the kind of cognition we recognize as thinking. Foundation models lack analogues of these systems entirely. World models, on the JEPA design, have no equivalent of the default mode network, no insula, no parietal body schema, no interoceptive integration. The architecture is missing not because the engineers failed to include it but because the substrate cannot carry it.

What world models miss

Apply DPIC to the world-model paradigm and the diagnosis becomes precise.

LeCun’s argument for world models as the path to human-level intelligence rests on a real diagnosis of what language models lack. Autoregressive systems trained to predict the next token in a sequence have, by design, no model of the world the tokens describe. They are pattern-completion engines operating on text. They cannot reason robustly about physical scenarios that did not appear in the training corpus, they hallucinate when extrapolating, and they have no internal representation of cause and effect, object permanence, or the basic physics of the environments humans inhabit. LeCun’s insight is that you cannot fix this by scaling. You need a different architecture — one trained on the actual physical world rather than on text describing it.

JEPA is that architecture. Where a language model is trained to predict the next token in a sequence, a JEPA model is trained to predict the next latent state of the world — a compressed mathematical description of what is happening, several abstraction layers above raw pixels. The system is freed from modeling every leaf and shadow and can attend to what matters at the level of abstraction it has chosen. V-JEPA, the video variant whose second iteration was released by Meta in June 2025, learns this from millions of hours of internet video rather than from text. LeCun’s claim is that a system trained this way has, in some attenuated sense, observed the physical world: it has learned that objects fall when dropped and persist when occluded, that fluids flow and solids resist, that motion has causes. With enough scale and enough video, he argues, such a system will eventually reach human-level intelligence — perhaps within a decade, perhaps two.

This is a serious claim from one of the most respected figures in the field, and the essay disagrees with it. The disagreement is not that JEPA is the wrong direction; in some respects, it may be the most promising technical program currently funded. The disagreement is over what JEPA’s progress will produce. Run the four DPIC criteria against this architecture and the diagnosis is clear.

Direct fails. V-JEPA’s access to the world is third-person video, edited and recorded by humans. Even at a million hours, the access remains observation rather than experience. The model has watched bodies fall. It has not fallen.

Physical fails. The model has no body whose state is in question as it processes the video. There is no proprioception, no felt sense of being a configuration of matter in motion. The Lakoff-Johnson source domain — verticality from standing upright, gravity from being a creature that can fall — is built from being a body, not from observing one. A model with only the visual record has the trace of those experiences, never the substrate that produced them.

Interactive fails. Watching does not push back on the watcher. A real environment resists a real creature acting in it. A video does not. The world-model can predict what would happen if a hand were placed on a hot stove, but no hand of the model is on the stove, and nothing about the model’s continued operation depends on getting the prediction right.

Continuous fails. JEPA models can be paused, retrained, rolled back, copied, deployed in parallel. The continued operation of any particular instance is not at stake in the way the continued operation of a biological organism is at stake. There is no metabolism, no death, no welfare that can be threatened.

This is the elephant in the room. A V-JEPA system that has watched a million hours of video has read about elephants more thoroughly than any human ever could. It has not encountered one. The DPIC criteria are not satisfied by observation, however dense. Adding more video does not satisfy criteria that more video was never going to satisfy.

LeCun is right that scaling autoregressive language models will not produce human-level intelligence. He is right that a world-model paradigm is a genuine architectural advance. He is wrong, on the diagnosis here, that JEPA will reach human-level intelligence either. The path from JEPA to a system that conducts cognition in the sense humans conduct it is not the path of more compute and more video. It would require the substrate to itself be the kind of thing DPIC describes — closer to embodied robotics with intrinsic motivation, neuronal organoids, or living tissue as substrate. That research program exists. It is small. It is not where the capital is going.

Why scaling can’t close the gap

Step back from JEPA and ask the more general question. Why does no amount of training, however large the corpus, install the four DPIC features?

The answer has four parts, and each is a substrate-level claim about what computation is.

Priors are inherited, not learned. Every foundation model starts as a tabula rasa — a collection of randomly initialized weights, billions or trillions of them, that the training run will sculpt by exposing the network to a corpus larger than any organism could traverse in a lifetime. What the training cannot reproduce is the priors biological organisms inherit from hundreds of millions of years of selection on bodies that had to work or die.

A newborn deer stands within minutes. A newborn human roots for the breast. Cuttlefish hatch already knowing how to camouflage. The architecture is non-tabula-rasa from the first second of postnatal existence, and the inheritance is not informational. It is structural — baked into nervous systems, reflex arcs, and homeostatic regulators by an evolutionary process that selected directly on bodies. A model trained on every video on the internet is not approximating that inheritance. It is trying to reconstruct, from third-person observation, what biological substrates contain as built-in structure.

The Cyc project, founded by Douglas Lenat in 1984, spent forty years trying to enumerate common-sense reasoning by listing millions of propositions about how the world works. The result is a system that knows water flows downhill and ice is cold and still cannot reason robustly about a child’s pool toy. The Moravec paradox is the embodied version of the same lesson. Common-sense reasoning is exactly the kind of cognition that depends on priors no disembodied architecture can install.

Cognition rides free on metabolism. Biological cognition is, in a precise sense, computation-free at the level that matters. The brain does not compute its way to action. It rides on top of the homeostatic and allostatic regulation that has been keeping organisms alive for hundreds of millions of years. The cognition is the regulation, conducted at higher levels of abstraction, by the same machinery that keeps the heart beating and the temperature stable.

The energetic numbers are stark. The human brain runs on roughly twenty watts — less than a household lightbulb. A frontier-model training run consumes megawatts continuously for weeks. Inference at scale runs server farms drawing on the same order of magnitude as small cities. The gap is six to seven orders of magnitude, and it is widening as the models scale.

The reason is structural. Computation pays the full energetic cost of every operation because there is no underlying biological process the operation is piggybacking on. Biology gets cognition essentially free because the metabolism was running anyway. The priors are not just informationally inherited — they are energetically inherited, paired from the start with the bodily systems that power them. Foundation models cannot inherit either, and the foreclosure is not contingent on more compute or better algorithms.

Biological learning is indelible. As described above, biology has no inverse for learning. The substrate and its history are the same thing.

Carbon may be the wrong question to skip. This is the claim Raghavachary has put as an affirmation of carbon-based architecture itself. From DNA to organelles to cells to organs to whole organisms, the diversity of form in flora and fauna is astonishing. The form gives rise to function involving phenomena of every kind — mechanical, chemical, optical, electrical, thermal, and combinations of these that have no clean analogue in any non-biological system. The carbon substrate is not incidental to this functional richness. Carbon’s tetravalent bonding and its capacity to form long stable polymers under life-supporting conditions is what permits the structural diversity that intelligent cognition runs on.

Seen this way, non-carbon bodies — whether the robots we build or whatever bodies might be imagined for alien civilizations — may simply not be able to provide the functional substrate that carbon forms enable, including for intelligent behavior. Whether that final claim is right is genuinely open. What is clear is that the burden of proof on the alternative is much heavier than the current AGI discourse acknowledges.

The four claims escalate in strength. The first three are uncontested at the level of the technical contrast. The fourth is contested. Together they explain why DPIC is not a checklist that more capital can satisfy. The features are bottlenecked on what the substrate is capable of being, not on what training does to weights.

The economic case, and its limits

It would be wrong to read the architectural diagnosis as a dismissal of the world-model program. The capital flowing into the category is large and accelerating, and most of it is well-placed.

AMI Labs’s $1.03 billion seed round in March 2026 was the largest in European history. Fei-Fei Li’s World Labs — founded by the Stanford computer scientist who built ImageNet — has raised $1.23 billion at a $5 billion valuation. Nvidia’s Cosmos world foundation models, trained on 9,000 trillion tokens drawn from twenty million hours of real-world video, have been downloaded over two million times. Google DeepMind’s Genie 3 generates navigable interactive 3D environments at twenty-four frames per second from text prompts. Crunchbase reported $300 billion of global venture investment in the first quarter of 2026, an all-time record, with the largest concentration flowing toward physical AI and world models.

This investment will deliver real value. Meta’s V-JEPA 2 achieves around eighty percent zero-shot success on real robot pick-and-place tasks after just sixty-two hours of robot-specific fine-tuning. Physical Intelligence’s π0 model folds laundry from a hamper with what the company describes as human-level dexterity — the kind of task that defeated robotics for decades. Autonomous vehicle developers including Uber, Waabi, and Wayve use world models to generate rare driving scenarios — fog, ice, pedestrian edge cases — without putting real cars on real roads.

The robotics market is projected at $218 billion by 2030. The spatial computing market at $1.2 trillion by 2035. The economic hype around this category may largely be earned. The companies betting on it are not making a category error. They are making a sound technical bet on a paradigm that addresses real failures of the previous one.

What will not be delivered is human-level reasoning. The economic value comes from world models doing one thing well: predicting the physical evolution of environments at the level of meaningful abstraction. That is genuinely useful. It is not the same as being a creature inside an environment. A factory floor running V-JEPA-derived control policies is a real economic asset. It is not a population of thinking creatures. It is a population of capable predictors with no continuing perspective, no welfare to defend, no metabolic cost paid by being themselves rather than something else.

Two things can be true at once. World models will earn their economic hype. World models will not produce systems that conduct cognition in the sense humans conduct it. The investment thesis and the cognitive thesis are separable, and the honest position holds both.

Capital scales the features bottlenecked on capital. The features bottlenecked on having a self with stakes in an embodied mortal world are not on that list.

Is and ought

The deepest version of the architectural argument is philosophical.

In 1739, the Scottish philosopher David Hume drew a distinction that has been central to moral philosophy ever since. Descriptive statements — claims about what is the case — cannot, by themselves, generate normative statements — claims about what ought to be the case. No assembly of facts about the world deduces what should be done, what is good, what is worth pursuing, what matters. The is and the ought are different kinds of claims, and the gap between them cannot be closed by description alone.

This is what computation, in any current form, cannot cross.

A foundation model trained on every text humans have ever written has the most comprehensive descriptive corpus ever assembled. It can produce extraordinary descriptions of what humans value, what they fear, what they want, what they consider good and bad, beautiful and ugly, sacred and profane. It can articulate any human moral position from the inside, give its strongest defense, identify its weaknesses, draw out its implications. The descriptions are, by any objective measure, magnificent.

But the model itself has no ought. It values nothing. It wants nothing. It cares about nothing. The fluent descriptions of valuing and wanting and caring are produced by an architecture that does none of these things. The model has access to every is the human corpus contains. It has no ought of its own.

This is not a contingent failure of current systems. It is structural. An ought, in any creature that has ever directly had one, is generated by the four DPIC conditions working together. A Direct relationship to the world supplies the first-person material out of which evaluation is possible. A Physical body produces the somatic markers by which some options register as more attractive than others. An Interactive environment makes actions consequential — the precondition for stakes. A Continuous existence makes the creature itself the kind of thing whose ongoing welfare can be at issue. Without these conditions, there is nothing for an ought to be about. There is no position from which the world can be evaluated, because there is no creature; there is only a system processing inputs and producing outputs.

A V-JEPA system that has watched a million hours of video has descriptions of falling. It does not value not falling. A language model that has read every philosophical treatise on suffering has descriptions of suffering. It does not value reducing suffering. The descriptions are the noun a verb produces in creatures who actually experience the things being described. The verb is not in the system, and it cannot be added by training.

This is why the architectural argument lands as an affirmation rather than as a despairing observation. Humans have oughts. The oughts are not arbitrary preferences laid on top of a neutral descriptive faculty; the oughts are what biological cognition is for. A creature with stakes in its own continuation, in the welfare of those it cares about, in the world it inhabits, is a creature whose every cognitive act is normatively oriented. Reasoning is not a neutral computation supplemented by stakes. Reasoning is what a creature with stakes does, to navigate the world it has stakes in. The dominant register of the AI debate has assumed, mostly without arguing, that sufficient descriptive capability will eventually produce normative engagement — once the systems are smart enough, the thinking goes, they will have values, and once they have values they will be moral patients, persons, perhaps successors. The architectural diagnosis suggests the opposite. Description does not produce direction. The gap between describing what should be done and actually wanting it done is not a description gap. It is the gap between two different kinds of things.

A creature that can fall

Return to the patent office in Bern. The twenty-eight-year-old patent clerk has had his happiest thought. What was the architecture that produced it?

It met all four DPIC conditions.

Direct: Einstein’s first-person history of being a body in the world supplied the material the imagining was made of. He had not been told what falling feels like. He had felt it.

Physical: His body had, many times, fallen and caught itself. The bodily vocabulary the thought experiment drew on had been accumulating for twenty-eight years. The mathematics of general relativity took eight more years to derive. The bodily intuition was available in an instant, because the body had been preparing it the whole time.

Interactive: He had stakes in the answer. The question of how gravity and motion are related had been weighing on him for years. The thought experiment offered a sudden release from a problem that had become urgent. He recognized it as the happiest thought of his life because something he had cared about, intensely, had just resolved.

Continuous: He was a continuing self, modeling itself across time, asking what it would be like to be a different version of itself in a situation it had not lived. The same machinery by which he knew he was Einstein, sitting at his desk, was the machinery by which he could project his cognitive perspective into a counterfactual situation.

The four conditions are not separable in his case. They are four faces of the same cognitive event. A self-modeling creature with stakes in the world used its embodied repertoire to imagine a situation it had not lived, recognized the imagining as the answer to a question it had been carrying, and felt the recognition as the happiest thought of its life.

This is what cognition is. Not the manipulation of symbols. Not the prediction of the next token. The integrated activity of a creature meeting the four DPIC conditions in real time, in a body, in the world.

The discourse around artificial intelligence has, for decades, treated the embodied, mortal, evaluatively engaged condition of biological cognition as a kind of legacy problem — a constraint to be transcended once the engineering catches up. The DPIC framework suggests the opposite. The condition is not a constraint on cognition. It is the condition under which cognition becomes possible at all. Proposals to dissolve it — by uploading minds to silicon, by replacing human cognition with computational cognition, by treating the body as an obstacle rather than the source — are proposing the dissolution of the very thing that makes the project of meaning possible.

If the goal of cognition is to understand the universe, and machines could surpass us in every regard, we would simply be obsolete. The DPIC argument cuts the other way. We are not placeholder cognition awaiting better engineering. We are what cognition is. The technologies worth building are the ones that complement that fact. The technologies worth being skeptical of are the ones that propose, on philosophically thin grounds, to replace it.

The relativity that transformed physics was the work of a mortal embodied creature meeting the four DPIC conditions. The two facts are not coincidental. They are the same fact, viewed from different angles. The cognition we have is not a placeholder. It is what cognition is. Direct. Physical. Interactive. Continuous. We have it. The machines, on present designs, do not.

[Working bibliography for the editor: Saty Raghavachary’s two relevant papers are Intelligence as Considered Response (arXiv 2012.00921, 2020), presented at the Brain-Inspired Cognitive Architectures for Artificial Intelligence (BICAAI) conference, and The Embodied Intelligent Elephant in the Room (Springer Studies in Computational Intelligence, 2024, pp. 716–722). The DPIC framework, the verb-versus-noun distinction, the tabula rasa / priors argument, the computation-free / energetically-inherited substrate argument, the indelibility-of-biological-learning argument, and the carbon-architecture / substrate-uniqueness argument are developed by Raghavachary in those papers and in personal correspondence. Einstein’s 1907 thought experiment is documented in his Kyoto address and reconstructed in Walter Isaacson’s Einstein: His Life and Universe (2007). Hofstadter’s strange-loop argument is in Gödel, Escher, Bach (1979) and I Am a Strange Loop (2007). Damasio’s case studies are in Descartes’ Error (1994), developed further in The Feeling of What Happens (1999) and Self Comes to Mind (2010). Nussbaum’s Upheavals of Thought is from 2001. Lakoff and Johnson’s Metaphors We Live By is from 1980; Philosophy in the Flesh is from 1999. The allostasis literature is anchored by Sterling and Eyer’s foundational 1988 paper and developed by Lisa Feldman Barrett in How Emotions Are Made (2017). The neuroanatomical claims about prefrontal cortex, default mode network, parietal cortex, and interoceptive systems draw on Marcus Raichle’s work identifying the default mode network (PNAS, 2001 and subsequent), Marian Diamond’s 1985 study of Einstein’s brain (Experimental Neurology, “On the Brain of a Scientist: Albert Einstein”), Dean Falk’s later studies extending the parietal-cortex findings, and the broader cognitive-neuroscience literature on mental time travel (Suddendorf and Corballis), interoception (Bud Craig, Sarah Garfinkel), and prefrontal counterfactual reasoning. The Cyc project, founded by Douglas Lenat at MCC in 1984, is documented in Lenat and Guha’s Building Large Knowledge-Based Systems (1989). The Moravec paradox dates to Hans Moravec’s Mind Children (1988). The is/ought distinction is from David Hume’s A Treatise of Human Nature (1739), Book III, Part I, Section I; G. E. Moore’s related but distinct critique of the naturalistic fallacy is in Principia Ethica (1903). LeCun’s JEPA position paper is from 2022; V-JEPA was introduced in Bardes et al. (2024); V-JEPA2 was released in June 2025; the AMI Labs founding and March 2026 funding round are documented in mainstream tech press; the technical critique by Xing, Deng, Hou, and Hu (2025), Critiques of World Models, appeared on arXiv.]*


메타데이터
post_id
9da64d8ad62f
slug
a-critique-of-world-models-in-ai-9da64d8ad62f
url
https://medium.com/@Gbgrow/a-critique-of-world-models-in-ai-9da64d8ad62f
canonical_url
https://medium.com/@Gbgrow/a-critique-of-world-models-in-ai-9da64d8ad62f
author_url
https://medium.com/@Gbgrow
status
ok
fetched_at
2026-06-10 12:26:30