Where Does Intelligence Actually Live? | The Substrate of Intelligence Series
An multipart inquiry into what thought is made of, the oldest question we ask about the mind, and why it has come back to decide the shape…
Where Does Intelligence Actually Live? | The Substrate of Intelligence Series
An multipart inquiry into what thought is made of, the oldest question we ask about the mind, and why it has come back to decide the shape of machine intelligence. Along the way: language, brain architecture, world models, and what human cognition teaches us about building and governing AI.
tl;dr — Almost everyone believes they think in words. A Turing laureate, a lab at MIT, patients who lost sight or speech yet kept reasoning intact, and a small model that learned Othello from nothing but move lists all suggest that belief is wrong. This is the premise of an eleven-part inquiry into what thought is made of: whether language is the seat of intelligence or one channel into it, and why that ancient question now decides how the largest AI systems get built. Two teams are spending billions on opposite bets, one training on language at scale, the other on video, action, and simulation. The position I defend across the series is that intelligence runs on a more fungible substrate than either camp assumes, and that compositional structure, more than any particular medium, is what decides whether something genuinely thinks.

Yann LeCun, a Turing laureate and one of the people who built modern deep learning, wrote that he does not think in words, and that animals do not think in words either. Pinned beneath the claim was a study from Evelina Fedorenko’s lab at MIT with a title that reads like a verdict: evidence from formal logical reasoning that the language of thought is not natural language [1]. Elsewhere he put the position plainly. Language may extend what we understand of the world. It does not exhaust intelligence. Corvids, octopuses, and primates, he said, are the proof.
Evidence from formal logical reasoning reveals that the language of thought is not natural language. Hope Kean, Alexander Fung, Paris Jaggers, Jason Chen, Joshua S. Rule, , Yael Benn, Joshua B. Tenenbaum, Steven T. Piantadosi5, Rosemary A. Varley, Evelina Fedorenko MIT, University College London, Manchester Metropolitan University, Cambridge University 2026 https://www.pnas.org/doi/10.1073/pnas.2520095123... joshrule.com/files/kean2025evidence.pdf
Here is the fact I was sure of. I think in words. Or I am nearly certain I do. There is a voice in my head narrating this sentence as I set it down, and when I work through a hard problem it mostly sounds like an argument I am having with myself. That voice feels so much like the seat of my thinking that for most of my life it never occurred to me to doubt it. Almost no one doubts it. This inner monlogue used to be my native language of Urdu, and for a long time now it has somehow switched to English. The belief that thought is essentially verbal is one of the quietest and most universal convictions we hold about our own minds, and we hold it the way we hold the floor beneath our feet, without ever noticing it is there.
And here was a serious mind in artificial intelligence, with a serious result behind him, telling me the floor was not where I was standing.

The disagreement does not stay personal for long. It is the same disagreement two well-funded teams are currently settling with hardware. Two teams can spend billions of dollars pursuing the same goal and build almost opposite machines. One team trains systems on language at immense scale. Its bet is that prediction becomes compression, compression recovers structure, and enough recovered structure becomes a usable model of the world. The other team trains on video, action, and simulation. Its bet is that intelligence begins with objects moving through space, causes producing effects, and agents discovering what happens when reality pushes back.
I have argued for architectures on both sides of that divide. I have sat with executives while a philosophical position became a platform decision, a compute commitment, and a multiyear budget.
The language of the discussion is technical. The disagreement underneath it is ancient: what is thought made of?
I could not let it go. This series is what came of that.

The suspicion under the certainty
The thing I want to name at the outset, because it took me a while to see it, is that my confidence proves less than it feels like it proves.
When I introspect, I find words. From that I concluded that thinking is done in words. But introspection reports the surface of experience, the part that reaches awareness, and there is no law that says the surface is where the work happens. My certainty is excellent evidence about what my thinking feels like. It is weak evidence about the machinery that produces the thought. The felt narration could be the computation itself, or it could be a readout of a computation that already finished somewhere I cannot see, arriving in words the way a printed receipt arrives after the transaction is settled.
That gap, between the form thought takes in awareness and the form it takes in the substrate, is the crack the whole series enters through. It is a genuinely philosophical gap and not a technical detail, because it puts a limit on the one instrument we most trust to tell us how we think. If the seat of reasoning is not open to introspection, then the question of what thought is made of cannot be answered by asking ourselves. It has to be answered from the outside, by evidence that does not care what our thinking feels like.
The people I kept running into made the gap impossible to ignore. Some report almost no inner speech. Some cannot form a mental image at will. Infants reason before they can speak. Animals plan without language. Adults keep reasoning after a stroke takes their grammar. My confidence in the voice was evidence about my experience. It was weak evidence about the machinery producing the thought. Three of those cases unsettled me enough to build the series around them.

Three cases where the essential channel disappears
Consider a person blind from birth. Vision is the channel through which most of us first meet space, shape, distance, and quantity, and yet a congenitally blind person develops the full apparatus of abstract and spatial reasoning, including advanced mathematics. More strikingly, the visual cortex, tissue we label by the sense it was meant to serve, gets recruited to process language and equations instead [2]. The channel we assumed was doing the work was never necessary. The mind routed around its absence and repurposed the very hardware that was supposed to be about seeing.
Consider a person whose stroke destroys the language network, leaving grammar in ruins. The intuition most of us carry is that reasoning rides on language, that to lose the words is to lose the thought. It does not happen that way. Patients with severe agrammatism can still perform exact multi-step calculation, follow causal chains, and reason about the world with the grammar gone [3]. Something is doing the reasoning that is not the thing making the sentences.
Now consider a machine. Take a transformer trained on nothing but sequences of legal Othello moves, a stream of symbols with no board attached. Probe its internal states and you find a representation of the board itself, a state that stands for the board, a live model of the game’s spatial structure that nobody ever supplied as input [4]. Intervene on that internal state and the model’s predictions change accordingly, which means the representation is being used, not merely present [5]. A system fed only a shadow of the game reconstructed the thing casting the shadow.
Three cases, three different sciences, three different kinds of mind. They share one unsettling feature. In each of them the channel that seems indispensable can vanish, and much of the intelligence remains. Vision goes missing and abstraction survives. Grammar goes missing and calculation survives. The board is never shown and the model of it appears anyway. Whatever intelligence is made of, it is not welded to the channel we were sure it needed.

An argument two thousand years older than the technology
The reason this grabbed me the way it did is that it is not a new question wearing new clothes. It is one of the oldest questions there is, and the people arguing about large language models today are, mostly without knowing it, re-fighting a battle that was fully formed long before anyone trained a network.
Aristotle, in On Interpretation, separated the spoken and written marks of language from the affections of the soul they signify, and observed that those affections are the same for everyone even though the words differ from tongue to tongue [6]. He was drawing a line between a public symbol and the private thing it points at, and that line runs straight through the present dispute. Medieval logicians posited a mental language beneath all spoken ones. Leibniz dreamed of a calculus of reason, a symbolic medium in which thinking would be computation. Against that whole lineage, a Romantic tradition running through Herder and Humboldt insisted that language does not dress a thought already complete, that it forms the thought, that a people’s tongue shapes the very world it can conceive.

Two camps, then, ancient and irreconcilable. One holds that structured thought precedes any particular language and merely borrows words to travel. The other holds that language is constitutive, that without it the higher reaches of thought never assemble at all. When two frontier labs pour billions of dollars of compute into opposite architectures, one betting that predicting enough text recovers a working model of the world, the other betting that text is a thin shadow and that a mind has to watch the world move before it can reason about anything real, they are taking sides in Aristotle’s argument with a budget attached. I find that both humbling and clarifying. The question is deep enough to have outlived every previous attempt to close it, which is a good reason to suspect it will outlive the current one too, unless we get the terms right.
Getting the terms right is most of the work, so let me spend a moment on the words the debate keeps sliding between.

The words the debate blurs
Three ideas get used interchangeably in casual talk, and pulling them apart dissolves a surprising amount of the confusion.
A representation is an internal state that stands in for something else so a system can operate on the stand-in instead of the world. A pattern of firing across neurons stands in for a remembered face. A vector in a model’s activation space stands in for a concept. The point is leverage. You manipulate the proxy, and if the representation is any good the manipulation tracks something true about the thing it represents.
An encoding is the format that carries the representation. The same content can wear different formats. The fact that the library sits northeast of the river can live in a map, in a sentence, in a sequence of remembered turns and step counts felt underfoot. Three formats, one relation, and a person who holds the relation in one format can usually rebuild it in another. That portability is the quiet miracle the whole series turns on.
A modality is the sensory or motor channel through which information arrives or departs: vision, hearing, touch, action, and the symbolic channel of language layered over them. Modalities are how a world gets into a mind and how a mind’s conclusions get back out.
And beneath all three sits the word I have put in the title. A substrate is the organization that makes representation and inference possible in the first place. It can mean physical implementation, neurons or silicon, or it can mean computational organization, the symbols, distributed vectors, variables, and rules over which thinking runs. Most of the noise in this field comes from collapsing those two senses, from arguing about the medium when the disagreement is really about the architecture, or the reverse.
With the words separated, the oldest live dispute in the science comes into focus, and it is a dispute about whether the format of a thought keeps the fingerprint of the channel it came in through.

Amodal cores and the grounding they owe
On one side is the idea of an amodal core.
On this view, when a mind holds a thought, the thought has been stripped of its sensory origin, the way a logical proof carries no residue of the eyes that first read the quantities. A representation is amodal when its computational role is independent of any single sensory or linguistic channel, and an amodal computational core is a system of such structured representations that receives information from several channels and supports reasoning across them.
Jerry Fodor gave this its strongest classical form. Cognition, he argued, runs over an internal, combinatorial symbol system, often called mentalese, that operates independently of any spoken language [7]. Its power rests on three properties worth naming, because they return in every part of this series. Compositionality: complex thoughts are built from reusable parts and the relations among them. Systematicity: a mind that can think “the dog chased the cat” can thereby think “the cat chased the dog,” because the same constituents recombine under the same rules, and you never meet a person who grasps the first and finds the second literally unthinkable. Productivity: from a finite stock of primitives, a mind generates an unbounded range of thoughts it has never met before. Alongside these, such a system needs a domain-general inference engine, a mechanism that can compare, predict, plan, and draw conclusions across several content domains. Fodor and Zenon Pylyshyn argued that these properties fall out of a structured architecture and have to be laboriously faked by one that lacks it [8]. On their telling, structure is the thing, and structure does not care whether its pieces arrived as sounds, sights, or signs.
The version I will defend is less doctrinaire than Fodor’s about implementation. The core could run on discrete symbols, distributed vectors, probabilistic programs, neural population codes, or an architecture nobody has named yet. The claim is about functional structure, not about a particular medium.
I am drawn to that view. I also think it owes a debt it cannot wave away, and the strongest objection in philosophy of mind is the one that names the debt.

The rival camp, embodied or grounded cognition, denies that a format-free core exists at all. Concepts, on this account, are partial reenactments of perception and action, and abstract thought is a faint re-running of physical experience. Lawrence Barsalou gave the position its influential cognitive form [9]. Behind the empirical version stands an older and harder philosophical one. Stevan Harnad’s symbol grounding problem asks how an internal symbol could mean anything at all, given that a symbol defined only by its relations to other symbols is a closed loop with no purchase on the world [10]. Emily Bender and Alexander Koller sharpened the same knife for the age of language models with a thought experiment: a system that has only ever seen the form of language, and never the world the language is about, has no path from form to meaning [11]. This is the deepest challenge the amodal view faces. Any account that puts a channel-independent core at the center of thought has to explain how the states in that core come to be about anything, how they earn their content rather than merely shuffling tokens. I take the objection as seriously as I take anything in this field, and I will keep returning to it, because I think a version of it survives even after the rest of the grounded position gives way.

The position I am prepared to defend, I think
Here is where I have landed for now, stated plainly enough that you can disagree with it before I have earned it.
Intelligence depends heavily on structured architecture and on a sufficiently rich model of a world. It depends far less on which channel supplies that structure. Language, vision, touch, action, and social contact are partially interchangeable routes into the same core. Their bandwidths differ, their biases differ, and their contributions are unequal across domains. Their identities do not settle what thinking is.
That view carries four convictions, and I will defend each across the eleven parts. Here, cognitive technology means a culturally created symbol system or practice that extends memory and reasoning:
- The conscious form of thought varies enormously from one person to the next, and it gives us a poor view of the computation underneath. What thinking feels like is close to useless as testimony about what thinking is.
- Language is a formidable cognitive technology, a communication system, an interface, an external memory, and in a bounded set of cases an instrument that helps construct concepts we could not otherwise hold. It is a tool the mind uses, and a magnificent one.
- Intelligence requires compositional structure and a sufficiently rich world model, an internal representation of entities, relations, causes, and dynamics that supports prediction and action.
- Grounding is the highest-bandwidth source of many physical and causal regularities, and language carries far more recoverable structure about the world than the strongest grounding critics allow.
The resulting position is an argument for architectural pluralism with a firm center. If I had to compress the whole wager to three sentences, it would be these. Words matter. Worlds matter. Structure decides what a mind can do with either.

The two questions that hold the series together
Underneath every part run two questions, and I want to name them now so you can hold me to them by name later.
The first is whether the substrate matters. In its popular form it asks whether intelligence is tied to words, or to pictures, or to sensorimotor experience. My answer splits the question in two, because the word substrate has been doing two jobs at once. If it means the channel, the words or images or motor routines a thought is carried in, then it is largely fungible, one able to substitute for another while the relevant competence survives. Blindness removes vision and abstraction still forms. Aphasia removes grammar and calculation still runs. If substrate means the architecture, the format at the level of computation, then it matters enormously and has a definite shape. A mind needs representations that keep the identity of things, bind them to roles, recombine them, and support causal intervention, reasoning about what would happen if the world were pushed. The channel is negotiable. The architecture is not.

The second is whether there is a critical mass, some minimum organization below which flexible thought never forms. My answer has two levels. There is a floor, and it is made of structure rather than of any particular sense or symbol. Below a certain compositional richness a system stays narrow and brittle, a stock of memorized responses with no capacity to generalize to genuinely new cases. Above that floor, intelligence looks like a continuum, rising with the breadth, depth, integration, and temporal reach of the world model, with particular tasks snapping into view suddenly once several gradually acquired capacities finally become available together. A threshold in the architecture, a continuum in the richness.
These two questions have different fates, and separating them dissolves most of the contradictions that sent me here. That is the claim the whole series exists to earn.

The map of the argument
Eleven parts, one line of inquiry. I will work from primary sources, give the strongest case on each side before I take mine, and say so plainly wherever the evidence genuinely conflicts rather than smoothing it over.
+------+---------------------------+--------------------------+---------------------------+
| Part | Central question | Main evidence | Direction of the answer |
+------+---------------------------+--------------------------+---------------------------+
| 1 | What is the substrate? | concepts and levels | separate channel, format, |
| | | of explanation | experience, architecture |
+------+---------------------------+--------------------------+---------------------------+
| 2 | How is the field divided? | eleven schools on two | the real fault line is |
| | | conceptual axes | channel vs. structure |
+------+---------------------------+--------------------------+---------------------------+
| 3 | How old is the dispute? | philosophy, linguistics, | structured thought and |
| | | cognitive science | linguistic tools coexist |
+------+---------------------------+--------------------------+---------------------------+
| 4 | How much does language | infants, cultures, | a prelinguistic core with |
| | shape thought? | space, number, color | real linguistic shaping |
+------+---------------------------+--------------------------+---------------------------+
| 5 | Where does reasoning run? | fMRI, lesions, math, | language and reasoning |
| | | formal logic | use separable networks |
+------+---------------------------+--------------------------+---------------------------+
| 6 | Is any channel required? | blindness, aphasia, | developing minds route |
| | | language deprivation | around missing channels |
+------+---------------------------+--------------------------+---------------------------+
| 7 | Does mindreading need | false belief, sign | early tracking is silent; |
| | language? | language, development | explicit recursion uses |
| | | | linguistic scaffolding |
+------+---------------------------+--------------------------+---------------------------+
| 8 | Do we think in words or | inner speech, imagery, | conscious format is a |
| | pictures? | experience sampling | variable readout |
+------+---------------------------+--------------------------+---------------------------+
| 9 | How much world is enough? | animals, deprivation, | compositional floor plus |
| | | scaling and emergence | a continuum of richness |
+------+---------------------------+--------------------------+---------------------------+
| 10 | Can text produce a mind? | LLMs, Othello, video, | text recovers structure; |
| | | world models | grounding fills key gaps |
+------+---------------------------+--------------------------+---------------------------+
| 11 | Where does the evidence | synthesis across fields | modality is fungible; |
| | finally land? | | structure is decisive |
+------+---------------------------+--------------------------+---------------------------+
The early parts clear the ground: the vocabulary, the map of schools, the two-thousand-year genealogy. The middle parts go where the evidence is hardest and least forgiving, into infants who reason before they can speak, brains that do logic with the language network idle, people blind or aphasic or raised without early language, and the sharpest test of all, whether reading another mind requires words. Part 8 is where my own starting premise falls apart, when it turns out that some people have no inner voice and some no mind’s eye and they think perfectly well anyway, which means the question I opened with, words or pictures, was broken from the start. The late parts turn to machines, to the scaling bet against the world-model bet, and to what all of this implies for anyone building these systems for real. Part 11 pays the debt and states the wager.
The route through the argument

Part 1: What intelligence is made of
Part 1 cleans up the vocabulary before taking a position. It separates representation, encoding, modality, conscious experience, external communication, computational architecture, and physical implementation. The familiar question, “Do you think in words or pictures?”, combines several of these into one and then treats the resulting confusion as a theory.

The first part also introduces phenomenology, the felt character of experience from the inside. Aphantasia, the absence of voluntary visual imagery, shows that spatial and visual competence can continue without a conscious picture [12]. Anendophasia, very low or possibly absent inner speech, has measurable effects on phonological tasks while leaving broad reasoning intact [13]. Descriptive experience sampling reports unsymbolized thinking, definite conscious thought with no experienced words or images [14]. These findings establish a limited point with large consequences: conscious format and computational format can come apart.
Part 1 gives both spine questions their formal shape. It also introduces the concession that governs the rest of the series. Phenomenological variation challenges the claim that conscious words or images constitute thought. It cannot, by itself, decide between subconscious sensorimotor simulation and amodal representation. That deeper question needs developmental, neural, and deprivation evidence.

Part 2: Eleven theories of how we think
Part 2 builds a map of the intellectual field. One axis runs from language-central accounts to language-peripheral accounts. The other runs from explanations of biological cognition to programs for machine intelligence. The map places eleven schools, including linguistic relativity, the view that a language shapes habitual cognition; core knowledge, early systems for objects, number, space, and agents; predictive processing, accounts that model cognition through prediction and error correction; the scaling thesis, the view that greater model, data, and compute scale can induce broad competence; and neuro-symbolic AI, systems that combine learned components with explicit structural or symbolic machinery. Language of thought, neural dissociation, generative linguistics, embodied cognition, world models, and deflationary accounts complete the map.

The map reveals that disciplinary labels often conceal the real disagreement. Fodor and the neuroscience of dissociation differ on the exact architecture, yet both place natural language outside the core machinery of general reasoning. Nativists and disciplined relativists can agree that infants possess prelinguistic structure while language later tunes habitual cognition. The deepest conflict runs diagonally across the map. Channel-centered theories place constitutive weight on language or sensorimotor experience. Structure-centered theories place it on internal organization that several channels can populate.
That diagonal also explains why current AI arguments feel strangely repetitive. Scaling advocates and world-model advocates appear to debate datasets and architectures. Their harder disagreement concerns whether a symbolic stream contains enough structure to induce a sufficiently rich model of reality.
Part 3: The idea of a language of thought
Part 3 develops the genealogy sketched earlier and carries it into the twentieth century. The ancient line from Aristotle through medieval mental language, Leibniz’s calculus of reason, and Fodor’s mentalese places structured thought prior to any particular public language [6]. The Romantic line through Herder and Humboldt, and later linguistic relativity, gives language a constitutive role in how a community organizes experience.

The twentieth century sharpens the dispute. Vygotsky argues that thought and speech begin separately, then converge as social speech becomes inner speech and a tool for self-direction [15]. Chomsky’s critique of Skinner helps end the behaviorist attempt to explain linguistic productivity through reinforcement alone and restores internal structure to scientific respectability [16]. Fodor and Pylyshyn then turn systematicity and productivity into architectural tests for connectionist models [8].
I land on a qualified thought-first view. The durable insight is that cognition operates over structured representations whose format exceeds any public language. Vygotsky preserves an equally durable insight: once acquired, language reorganizes parts of cognition, extends working memory, stabilizes abstractions, and enables recursive self-instruction. The exact internal format remains open. A tidy classical symbol system is one proposal among several.
Part 4: Does language shape how you think?
Part 4 puts nativism and linguistic relativity against the evidence. Research on core knowledge suggests that substantial structure is present before fluent language [17]. Susan Carey’s account of Quinian bootstrapping, where placeholder symbols help children construct concepts that exceed their starting systems, explains how culture can add genuine conceptual machinery [18].

Cross-linguistic research then shows language shaping habitual cognition. Speakers whose languages rely on absolute spatial coordinates can organize nonverbal spatial memory differently from speakers of relative-frame languages [19]. The Piraha number studies supply a sharper result. Speakers without exact number words can match quantities when memory demands stay low, yet exact large quantities become difficult to preserve across delay. Frank and colleagues describe number words as a cognitive technology for maintaining cardinality across time, space, and modality [20].
My conclusion gives both sides a real victory. A prelinguistic representational core is robust. Language tunes attention and memory, selects coordinate frames, and helps build exact integer concepts. That last role is constitutive and bounded. It prevents the easy claim that language always serves as a detachable label for concepts already complete.
Part 5: How the brain encodes a thought
Part 5 moves from behavior to tissue. A population code is information carried by the joint activity of many neurons. A multiple-demand network is a distributed frontoparietal system recruited for flexible control and difficult problem solving. The human language network occupies largely left frontal and temporal regions and can be localized separately from systems used for arithmetic, formal reasoning, music, code, and social inference.

Fedorenko, Piantadosi, and Gibson synthesize evidence for a double dissociation between language and thought, arguing that language primarily supports communication and cultural transmission [21]. Kean and colleagues report that formal logical reasoning recruits circuitry separable from the natural-language network [1]. Amalric and Dehaene find that advanced mathematics in professional mathematicians activates bilateral frontal and parietal regions while sparing language-related areas [22]. Lesion evidence converges from the other direction: severe grammar loss can coexist with exact calculation [3].
I treat this as the strongest evidence in the series that adult reasoning and natural-language processing are separable functions. The developmental question remains: a mature capacity can run independently after language helped build it. Parts 6 and 7 address that gap.
Part 6: Intelligence without sight or speech
Part 6 examines natural experiments created by congenital blindness, aphasia, and delayed access to language. Congenital blindness shows cortical pluripotency, the developing cortex’s capacity to take on different high-level functions depending on available input. Occipital tissue can participate in language and mathematics when visual experience never arrives [2]. Severe aphasia shows calculation and causal reasoning surviving profound grammatical damage [3]. Deaf children without early access to a shared language often show intact nonverbal intelligence alongside a targeted delay in explicit false-belief reasoning [23].

The conclusion requires precision. No individual sensory or linguistic channel is necessary for broad intelligence. Development finds alternate routes. This demonstrates redundancy and plasticity. It also shows that timing and structure matter. Some capacities depend on access to rich, socially shared representation during a sensitive period. The modality can vary, as fluent sign language demonstrates. The requirement concerns structured input arriving while the architecture is still developing.
Part 6 therefore strengthens both halves of my thesis. Channel identity is negotiable. Developmental conditions place genuine constraints on what the eventual system can represent.
Part 7: Do you need words to read a mind?
Part 7 gives the language-dependent case its strongest hearing through theory of mind, the ability to represent beliefs, desires, and intentions in other agents. A standard false-belief task requires a child to keep the real location of an object separate from another person’s mistaken belief about it. Explicit success usually appears around age four or five and correlates with mastery of sentential complements, constructions that embed one proposition inside another, such as “Sally thinks [that the marble is in the basket].”

Language deprivation and the emergence of Nicaraguan Sign Language make the relationship difficult to dismiss. Richer mental-state language predicts gains in explicit false-belief understanding [24]. Yet preverbal infants show earlier sensitivity to belief-like states in looking-time paradigms [25], and adult belief reasoning recruits circuitry separate from the language network. Apperly and Butterfill’s two-systems account reconciles the pattern: an early, fast system tracks limited belief-like states, while a later, flexible system supports explicit, recursive metarepresentation, the capacity to represent another representation as true or false [26].
Language helps build and operate the later system. This is a meaningful qualification to channel fungibility. It shows that a symbolic technology can deepen the architecture without becoming the universal medium of all thought.
Part 8: Do you think in words or pictures?
Part 8 returns to the introspective question that started the project. People report vivid inner speech, vivid imagery, mixed forms, sparse conscious content, and unsymbolized thought. Aphantasia leaves visual recognition, navigation, and much spatial reasoning available [12]. Low inner speech produces bounded costs in verbal working memory and rhyme judgment [13]. Lind’s methodological critique cautions against treating questionnaire reports as proof of a population with literally zero inner speech [27]. Hurlburt’s sampling work nevertheless supports wide variation in the experienced form of thought [14].

My conclusion is narrower than some amodal theorists would prefer. The medium of conscious thought behaves like an interface style. It changes the texture of mental life and supports certain tasks. It predicts little about general reasoning power. Introspection reveals the display available to a person, with limited access to the computation producing it.
This evidence cannot settle the embodied-cognition dispute. Subconscious sensorimotor simulation and amodal representation can both remain invisible to awareness. It does settle the naive words-or-pictures debate. Conscious verbal and visual formats are optional readouts.
Part 9: How much world does a mind need?
Part 9 confronts critical mass directly. Animal cognition offers natural points along a continuum of world-model richness. Grey parrots can classify objects, compare properties, and make quantity judgments [28]. Scrub jays bind what, where, and when in episodic-like memory and adjust caching behavior to future conditions [29]. Chimpanzees display tool use and, with symbolic training, forms of relational reasoning. These cases show structured internal models operating in different brains without human language.
Machine learning adds a dispute about thresholds. Reports of emergent abilities, capabilities that appear abruptly at a certain scale, suggest phase transitions [30]. Schaeffer and colleagues show that some abruptness comes from discontinuous metrics applied to smoothly improving underlying probabilities [31]. Both observations can hold. Representational capacity may accumulate gradually while a task becomes solvable only after several components reach a usable configuration.

I therefore propose a two-level answer. Compositional organization supplies a floor for flexible recombination. Above that floor, intelligence is continuous and tracks the richness of the world model. Raw parameter count is an input to that process, not its definition. The more relevant measures are structural depth, causal reach, memory, counterfactual flexibility, and generalization across novel combinations.
Part 10: Words versus worlds
Part 10 translates the entire series into the dispute shaping frontier AI. The scaling thesis holds that sufficiently demanding prediction over language forces a model to recover the latent causes that produced the text. Othello-GPT is an important proof of possibility. A model trained on move sequences develops an internal board representation that probes can recover and interventions can alter [4], [5]. A symbolic sequence can therefore induce a genuine domain model.
The world-model thesis argues that autonomous intelligence requires models learned from the dynamics of reality. LeCun’s energy-based and joint-embedding program predicts abstract representations of future states [32]. V-JEPA 2 extends this direction into understanding, prediction, and planning from video [33]. Ha and Schmidhuber’s earlier world-model architecture showed an agent learning inside a compressed latent simulation [34]. These programs focus on the information text systematically omits: fine spatial geometry, intuitive physics, persistence, affordances, and the consequences of action.

The deflationary challenge remains essential. Bender and Koller argue that linguistic form alone cannot supply communicative intent and grounded meaning [11]. The neuro-symbolic response seeks compositional causal models, intuitive physics, and intuitive psychology alongside learning [35].
My adjudication favors convergence. Text contains enough structured projection of the world to support powerful reconstruction. Grounded multimodal streams supply physical and causal information with far greater fidelity. A capable architecture should combine learned world models, compositional abstractions, domain-general inference, language interfaces, memory, and tools.
Part 11: Does the substrate really matter?
Part 11 answers both spine questions and states the wager. At the level of input channel and conscious format, the substrate is substantially fungible. At the level of computational architecture, it matters enormously. Flexible intelligence requires structured representations, systematic recombination, causal modeling, memory, and variable binding, the assignment of entities to roles while preserving their relations. The initial stock may include innate priors, built-in biases that direct learning toward objects, number, space, and agents, though the exact balance between learned structure and priors remains unsettled.

Critical mass is a compositional floor followed by a continuum. The floor concerns the ability to bind roles and relations and to generalize across new combinations. The continuum concerns how much world the system can represent, how far it can predict, how well it can revise itself, and how flexibly it can move across contexts.
My final wager favors systems that learn hierarchical world models from language and grounded experience, carry explicit or reliably emergent compositional structure, and use language as interface, instruction medium, coordination layer, and external memory. This wager can lose. Text may yield more physical structure than I expect. Priors may emerge from training with less explicit engineering. Embodied approaches may prove necessary in more domains. Those are empirical questions, and the architecture should be judged by intervention and generalization, independent of philosophical allegiance.

The practical throughline
The substrate question changes how I build, buy, and govern AI systems.
For builders, the first consequence is architectural. A fluent language model can serve as an excellent interface and a surprisingly capable reasoner. Numerical, logical, physical, and high-assurance tasks still benefit from explicit solvers, verifiers, tools, structured memory, and models of state. A robotics system needs grounded dynamics. A contract-analysis system can learn much of its world from text because the governing domain is largely encoded in documents. A clinical workflow needs language, structured patient state, temporal causality, and carefully bounded tools. The richest useful channel depends on the domain.
For buyers, the key consequence is evaluative. Fluency measures the interface. Reasoning quality requires separate tests. I would evaluate entity tracking, variable binding, compositional generalization, counterfactual consistency, causal intervention, tool use, recovery from changed premises, and performance on held-out combinations. I would also test whether the internal state remains coherent as the surface wording changes. A high benchmark average can hide a system that interpolates elegantly and breaks on a simple structural variation.
For governors, the consequence is epistemic humility. A model’s verbal explanation is an output generated through the same interface being evaluated. Chain of thought is a verbalized sequence that may help produce or explain an answer, yet it can also become incomplete, post-hoc, or strategically shaped. Assurance requires behavioral tests, controlled interventions, provenance, tool logs, and outcome monitoring. The narration can support audit. It cannot carry the audit by itself.
The critical-mass argument also changes capability monitoring. Pass-fail benchmarks can make a gradually rising capability appear sudden. Continuous measures can expose the slope before deployment crosses a practical threshold. Responsible governance should track both: gradual changes in underlying competence and discrete task boundaries that create operational risk.
The board-level question is therefore sharper than “Which model is smartest?” I would ask: What world does this system need to represent? Which channels contain that world’s relevant structure? Which parts require grounding, explicit state, or deterministic tools? Where must the system generalize compositionally? What evidence would show that it actually does?
Those questions turn philosophy into architecture, procurement criteria, and governance controls.

Where I land
My position is clear enough to challenge.
Intelligence resides in structured internal organization that can encode entities, roles, relations, causes, and alternatives, then operate over them flexibly. Natural language gives human minds an extraordinary instrument for communication, abstraction, memory, recursion, and cultural accumulation. Sensory experience gives those minds dense contact with physical and social reality. Each contributes capabilities the other supplies poorly.
The identity of the channel remains secondary to the quality of the structure acquired through it. The structure itself is indispensable. General machine intelligence will require compositional representations, rich causal world models, domain-general inference, memory, grounding appropriate to the target domain, and interfaces through which humans can teach, inspect, and govern the system.
I expect text scaling to keep producing substantial gains. I also expect diminishing returns in domains whose decisive regularities rarely appear in text. I would place capital behind hybrid systems that join language models, predictive world models, structured abstractions, tools, and verifiers. The winning design will treat words as a cognitive technology, grounded experience as a source of missing structure, and internal representation as the substrate that makes both useful.
Stated as a dispute, the shape is this. The naive claim dies. You do not have to think in words, you do not have to think in pictures, and the inner voice I trust so completely is an instrument my mind reaches for on certain jobs rather than the place the thinking happens. The fashionable opposite claim, that the substrate is irrelevant and only scale decides, is hiding something too. When you look closely at what survives every deprivation and every counterexample, two things remain standing. Some compositional structure to hold thoughts in. Some real contact with a world to make those thoughts about anything at all. Those are constraints on the substrate. They are simply not the constraint everyone is arguing about.
Several questions remain open: the exact neural and computational format of amodal representation, the minimum grounding required by different tasks, the right balance between learned structure and built-in priors, the reliability of mechanistic probes, and the relationship between computation and consciousness, subjective experience or the presence of something it feels like to be a system. The later parts examine these uncertainties without pretending to close them.
That is the wager. Over the next eleven parts I am going to try to earn it, or to find out I was wrong. Look past the voice, the image, and the polished answer. Ask what structure had to exist for the thought to become possible.
Come argue with me.

Frequently asked questions
1. Are you claiming that language is irrelevant to thought?
No. Language transforms cognition. It supports inner rehearsal, shared abstraction, cultural accumulation, exact number, explicit recursive belief reasoning, instruction, and external memory. The claim is narrower: broad reasoning can exist before language, outside the language network, and after severe grammatical loss. Language amplifies and reorganizes a thinking system with substantial prelinguistic structure.
2. Does aphantasia prove that concepts are amodal?
No. Aphantasia shows that conscious visual imagery can be absent while visual knowledge and much spatial reasoning remain. An embodied theorist can still propose subconscious sensorimotor simulation. Aphantasia separates conscious display from competence. Neural, lesion, developmental, and intervention evidence must carry the deeper argument about representational format.
3. Does congenital blindness refute embodied cognition?
It refutes versions that make visual experience necessary for abstract concepts or mathematics. A broader embodied account can appeal to touch, action, proprioception, language, and social interaction. Blindness demonstrates developmental plasticity and channel substitution. It leaves grounding through other forms of experience available.
4. Was Fodor right about the language of thought?
Fodor identified enduring architectural requirements: systematicity, productivity, and compositional structure. His claim that cognition requires an innate classical symbol system remains contested. Modern neural networks can learn distributed representations and sometimes develop compositional internal structure. Their generalization remains uneven. I retain Fodor’s tests while leaving his preferred implementation open.
5. Can a text-only model develop a genuine world model?
Yes, within bounded domains. Othello-GPT provides causal evidence that sequence prediction can induce an internal model of an unseen state space [4], [5]. The open question concerns breadth. Natural text is a selective projection of physical and social reality. It carries rich causal and cultural information while omitting much fine geometry, intuitive physics, and action feedback.
6. Does intelligence have a threshold?
Some architectural requirements behave like thresholds. A task requiring variable binding or several coordinated representations can fail until the full configuration is available. Underlying capacity can still grow smoothly. I describe this as a compositional floor followed by a continuum of world-model richness.
7. Is Othello-GPT too small and artificial to support a general claim?
It supports a possibility claim: a thin symbolic stream can induce structured latent state. It offers weak evidence about open-ended physical intelligence because Othello is closed, deterministic, fully specified, and small. The case defeats the universal claim that sensory grounding is the only route to internal structure. It leaves the sufficiency of text for broad intelligence unresolved.
8. Why should an enterprise care about this debate?
Because architecture, evaluation, and risk controls depend on it. Text-rich domains can rely heavily on language models. Physical and temporal domains need grounded state and causal models. High-assurance reasoning benefits from solvers and verifiers. Procurement should test structural generalization. Governance should use interventions and behavior alongside verbal explanations.
9. Does this series claim that current AI systems are conscious?
No. The series concerns representation, inference, learning, and behavior. Computational competence alone does not settle whether subjective experience exists. That question remains outside the argument developed here.
Glossary of key terms
Amodal: Independent of any single sensory or linguistic channel at the level relevant to computation.
Amodal computational core: A system of structured, channel-independent representations used for reasoning across domains.
Anendophasia: Very low or possibly absent inner speech, associated with bounded costs on some verbal tasks.
Aphantasia: The absence of voluntary visual imagery despite preserved visual knowledge and many visual or spatial abilities.
Chain of thought: A verbalized sequence presented as intermediate reasoning, which may be useful without being a complete causal record of computation.
Cognitive technology: A culturally created symbol system or practice that extends memory, reasoning, or communication.
Compositionality: The capacity to build complex representations from reusable parts and the relations among them.
Consciousness: Subjective experience, or the presence of something it feels like to be a system.
Core knowledge: Early-emerging representational systems for domains such as objects, number, space, and agents.
Cortical pluripotency: The developing cortex’s capacity to support different cognitive functions depending on input and connectivity.
Critical mass: The minimum organization or representational richness required for a particular cognitive capacity.
Domain-general inference engine: A mechanism that supports flexible reasoning, planning, comparison, and control across content domains.
Embodied cognition: The family of theories that gives bodily states, action, and sensorimotor systems a constitutive role in concepts and thought.
Encoding: The format in which information is carried inside or between systems.
False-belief task: A test of whether someone can represent another person’s belief when it conflicts with reality.
Fungibility: The degree to which one channel or format can substitute for another while preserving the relevant competence.
Grounded cognition: The view that representations gain content through connections to perception, action, and the world.
Innate priors: Built-in learning biases that guide a developing system toward useful structures such as objects, space, number, or agents.
Language of thought: The hypothesis that cognition operates over an internal, structured representational medium distinct from public language.
Linguistic relativity: The view that the language a person speaks shapes habitual patterns of attention, memory, or categorization.
Mentalese: Fodor’s name for the internal combinatorial representational language proposed to underlie thought.
Metarepresentation: A representation of another representation, such as understanding that a person holds a belief that can be false.
Modality: A channel through which information enters or leaves a system, such as vision, hearing, touch, action, or language.
Multiple-demand network: A frontoparietal brain network associated with flexible control, working memory, and difficult problem solving.
Neuro-symbolic AI: Approaches that combine learned neural representations with explicit symbols, rules, or structured abstractions.
Phenomenology: The felt character and conscious form of experience from the first-person point of view.
Population code: Information represented by a pattern of activity distributed across many neurons.
Predictive processing: A family of accounts that describes cognition through prediction, error correction, and model updating.
Productivity: The ability to generate an open range of new thoughts from finite representational resources.
Quinian bootstrapping: Carey’s proposal that children use learned placeholder symbols to construct concepts beyond their initial core systems.
Representation: An internal state that stands for something else and can be manipulated in its place.
Scaling thesis: The view that prediction over sufficiently large datasets and models can induce broad competence and latent world structure.
Sentential complement: A grammatical construction that embeds one proposition inside another, allowing the embedded proposition to be represented as true or false.
Substrate: The physical or computational organization that makes representation and inference possible.
Symbol grounding problem: The challenge of explaining how internal symbols acquire meaning beyond their relations to other symbols.
Systematicity: The capacity to entertain structurally related recombinations of a thought’s constituents.
Theory of mind: The ability to attribute and reason about beliefs, desires, intentions, and other mental states.
Two-systems account: The proposal that early, fast belief tracking and later, explicit belief reasoning rely on distinct cognitive systems.
Unsymbolized thinking: A reported conscious thought with definite content and no experienced words, images, or other symbols.
Variable binding: The assignment of an entity to a role in a structured representation while preserving the surrounding relations.
World model: An internal representation of entities, relations, causes, and dynamics that supports prediction, counterfactual reasoning, and action.
World-model thesis: The view that broad intelligence requires learning structured dynamics from grounded, temporally rich experience.
References
[1] H. H. Kean, A. Fung, R. P. Levy, and E. Fedorenko, “Evidence from formal logical reasoning reveals that the language of thought is not natural language,” Proceedings of the National Academy of Sciences, vol. 123, no. 28, Art. no. e2520095123, 2026. [Online]. Available: https://doi.org/10.1073/pnas.2520095123
[2] M. Bedny, “Evidence from blindness for a cognitively pluripotent cortex,” Trends in Cognitive Sciences, vol. 21, no. 9, pp. 637–648, 2017. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/28821345/
[3] R. A. Varley, N. J. C. Klessinger, C. A. J. Romanowski, and M. Siegal, “Agrammatic but numerate,” Proceedings of the National Academy of Sciences, vol. 102, no. 9, pp. 3519–3524, 2005. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/15713804/
[4] K. Li, A. K. Hopkins, D. Bau, F. Viegas, H. Pfister, and M. Wattenberg, “Emergent world representations: Exploring a sequence model trained on a synthetic task,” in International Conference on Learning Representations, 2023. [Online]. Available: https://arxiv.org/abs/2210.13382
[5] N. Nanda, A. Lee, and M. Wattenberg, “Emergent linear representations in world models of self-supervised sequence models,” in BlackboxNLP, 2023. [Online]. Available: https://arxiv.org/abs/2309.00941
[6] Aristotle, On Interpretation, E. M. Edghill, Trans. [Online]. Available: https://classics.mit.edu/Aristotle/interpretation.html
[7] J. A. Fodor, The Language of Thought. Cambridge, MA, USA: Harvard University Press, 1975. [Online]. Available: https://books.google.com/books/about/The_Language_of_Thought.html?id=XZwGLBYLbg4C
[8] J. A. Fodor and Z. W. Pylyshyn, “Connectionism and cognitive architecture: A critical analysis,” Cognition, vol. 28, nos. 1–2, pp. 3–71, 1988. [Online]. Available: https://www.sciencedirect.com/science/article/pii/0010027788900315
[9] L. W. Barsalou, “Perceptual symbol systems,” Behavioral and Brain Sciences, vol. 22, no. 4, pp. 577–660, 1999. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/11301525/
[10] S. Harnad, “The symbol grounding problem,” Physica D: Nonlinear Phenomena, vol. 42, nos. 1–3, pp. 335–346, 1990. [Online]. Available: https://www.cs.ox.ac.uk/activities/ieg/e-library/sources/harnad90_sgproblem.pdf
[11] E. M. Bender and A. Koller, “Climbing towards NLU: On meaning, form, and understanding in the age of data,” in Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020, pp. 5185–5198. [Online]. Available: https://aclanthology.org/2020.acl-main.463/
[12] A. Zeman, M. Dewar, and S. Della Sala, “Lives without imagery: Congenital aphantasia,” Cortex, vol. 73, pp. 378–380, 2015. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/26115582/
[13] J. S. K. Nedergaard and G. Lupyan, “Not everybody has an inner voice: Behavioral consequences of anendophasia,” Psychological Science, vol. 35, no. 7, pp. 780–797, 2024. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/38728320/
[14] R. T. Hurlburt and S. A. Akhter, “Unsymbolized thinking,” Consciousness and Cognition, vol. 17, no. 4, pp. 1364–1374, 2008. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/18456514/
[15] L. S. Vygotsky, Thought and Language, E. Hanfmann and G. Vakar, Trans. Cambridge, MA, USA: MIT Press, 1962. [Online]. Available: https://mitpress.mit.edu/9780262720014/thought-and-language/
[16] N. Chomsky, “A review of B. F. Skinner’s Verbal Behavior,” Language, vol. 35, no. 1, pp. 26–58, 1959. [Online]. Available: https://chomsky.info/1967____/
[17] E. S. Spelke and K. D. Kinzler, “Core knowledge,” Developmental Science, vol. 10, no. 1, pp. 89–96, 2007. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/17181705/
[18] S. Carey, The Origin of Concepts. New York, NY, USA: Oxford University Press, 2009. [Online]. Available: https://academic.oup.com/book/12686
[19] S. C. Levinson, S. Kita, D. B. M. Haun, and B. H. Rasch, “Returning the tables: Language affects spatial reasoning,” Cognition, vol. 84, no. 2, pp. 155–188, 2002. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/12175571/
[20] M. C. Frank, D. L. Everett, E. Fedorenko, and E. Gibson, “Number as a cognitive technology: Evidence from Piraha language and cognition,” Cognition, vol. 108, no. 3, pp. 819–824, 2008. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/18547557/
[21] E. Fedorenko, S. T. Piantadosi, and E. A. F. Gibson, “Language is primarily a tool for communication rather than thought,” Nature, vol. 630, pp. 575–586, 2024. [Online]. Available: https://doi.org/10.1038/s41586-024-07522-w
[22] M. Amalric and S. Dehaene, “Origins of the brain networks for advanced mathematics in expert mathematicians,” Proceedings of the National Academy of Sciences, vol. 113, no. 18, pp. 4909–4917, 2016. [Online]. Available: https://doi.org/10.1073/pnas.1603205113
[23] C. C. Peterson and M. Siegal, “Deafness, conversation and theory of mind,” Journal of Child Psychology and Psychiatry, vol. 36, no. 3, pp. 459–474, 1995. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/7782409/
[24] J. E. Pyers and A. Senghas, “Language promotes false-belief understanding: Evidence from learners of a new sign language,” Psychological Science, vol. 20, no. 7, pp. 805–812, 2009. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/19515119/
[25] K. H. Onishi and R. Baillargeon, “Do 15-month-old infants understand false beliefs?,” Science, vol. 308, no. 5719, pp. 255–258, 2005. [Online]. Available: https://www.science.org/doi/10.1126/science.1107621
[26] I. A. Apperly and S. A. Butterfill, “Do humans have two systems to track beliefs and belief-like states?,” Psychological Review, vol. 116, no. 4, pp. 953–970, 2009. [Online]. Available: https://www.butterfill.com/writing/two_systems/
[27] A. Lind, “Are there really people with no inner voice? Commentary on Nedergaard and Lupyan (2024),” Psychological Science, vol. 36, no. 9, pp. 765–767, 2025. [Online]. Available: https://pubmed.ncbi.nlm.nih.gov/40424755/
[28] I. M. Pepperberg, “Cognitive and communicative abilities of Grey parrots,” Applied Animal Behaviour Science, vol. 100, nos. 1–2, pp. 77–86, 2006. [Online]. Available: https://www.sciencedirect.com/science/article/abs/pii/S0168159106001055
[29] N. S. Clayton and A. Dickinson, “Episodic-like memory during cache recovery by scrub jays,” Nature, vol. 395, pp. 272–274, 1998. [Online]. Available: https://www.nature.com/articles/26216
[30] J. Wei et al., “Emergent abilities of large language models,” Transactions on Machine Learning Research, 2022. [Online]. Available: https://arxiv.org/abs/2206.07682
[31] R. Schaeffer, B. Miranda, and S. Koyejo, “Are emergent abilities of large language models a mirage?,” in Advances in Neural Information Processing Systems 36, 2023. [Online]. Available: https://arxiv.org/abs/2304.15004
[32] A. Dawid and Y. LeCun, “Introduction to latent variable energy-based models: A path towards autonomous machine intelligence,” arXiv:2306.02572, 2023. [Online]. Available: https://arxiv.org/abs/2306.02572
[33] M. Assran et al., “V-JEPA 2: Self-supervised video models enable understanding, prediction and planning,” arXiv:2506.09985, 2025. [Online]. Available: https://arxiv.org/abs/2506.09985
[34] D. Ha and J. Schmidhuber, “World models,” arXiv:1803.10122, 2018. [Online]. Available: https://arxiv.org/abs/1803.10122
[35] B. M. Lake, T. D. Ullman, J. B. Tenenbaum, and S. J. Gershman, “Building machines that learn and think like people,” Behavioral and Brain Sciences, vol. 40, Art. no. e253, 2017. [Online]. Available: https://arxiv.org/abs/1604.00289
메타데이터
- post_id
- 4f1944d455fe
- slug
- where-does-intelligence-actually-live-the-substrate-of-intelligence-series-4f1944d455fe
- url
- https://medium.com/@adnanmasood/where-does-intelligence-actually-live-the-substrate-of-intelligence-series-4f1944d455fe
- canonical_url
- https://medium.com/@adnanmasood/where-does-intelligence-actually-live-the-substrate-of-intelligence-series-4f1944d455fe
- author_url
- https://medium.com/@adnanmasood
- status
- ok
- fetched_at
- 2026-07-21 21:16:41