Knowledge buried in Language
What if the parrot isn’t dumb, but fluent? From Skinner’s behaviorism to Chomsky’s universal grammar, and from Wittgenstein’s language…
Knowledge buried in Language
What if the parrot isn’t dumb, but fluent? From Skinner’s behaviorism to Chomsky’s universal grammar, and from Wittgenstein’s language games to the rise of AI, we trace how language models unsettle what it means to understand, speak, and mean.

“I think this Skinner box will fix the errors” — made with OpenArt
The Parrot’s Revenge
When Joseph Weizenbaum built ELIZA in the 1960s, a simple chatbot mimicking a therapist, he was shaken by how quickly users opened up to it. Even his own secretary asked to be left alone with the machine. It was just pattern-matching and scripted responses, but people felt heard.
Weizenbaum warned that we too easily project meaning onto machines. The danger, he believed, was that we mistook the machine’s responses as representing understanding.
Now, the parrot is back. Language models like ChatGPT can respond convincingly to almost anything, without “knowing” anything at all. They learn exactly as B.F. Skinner once imagined: by observing past language use. And yet, they seem to know.
So the question arises: has behaviorism staged a comeback, or are we entering a new epistemic crisis, with AI playing the role of the verbal zombie?
Skinner’s behaviorism
B.F. Skinner saw language not as thought, but as action. In Verbal Behavior (1957), he proposed that children learn to speak through exposure, reinforcement, and repetition. A child says “water” and receives water: the behavior is reinforced. Language, in this model, is shaped entirely by external forces.
Thinking in and of itself, Skinner claimed, was nothing but internal verbal behavior. A kind of silent self-talk internalized through the same mechanisms that shape spoken language. The inner monologue had no privileged status. It was just soundless speech.
This view aligned well with mid-century psychology, which dismissed mental states as unscientific speculation. If it couldn’t be observed, it wasn’t worth studying. Language, in this light, was no more special than pushing a button or pulling a lever.
Chomsky’s Revolt
In 1959, Noam Chomsky published a scathing review of Verbal Behavior in the journal Language. He challenged the very premise of Skinner’s theory, namely that language could be explained without invoking internal structures or innate capacities.
Chomsky’s main point was that language is creative. We utter sentences we’ve never heard before, and understand sentences we’ve never been taught. No amount of conditioning can explain our ability to distinguish “Colorless green ideas sleep furiously” from “Furiously sleep ideas green colorless” without invoking underlying grammatical structure.
He pointed to examples like “Your money or your life,” where comprehension, he argued, cannot be the result of generalizing over past experience. We don’t learn its meaning from statistical patterns involving “money” and “life,” but from an internal grasp of what it means to live, to be threatened, to feel fear.
Chomsky proposed there were specific language learning mechanisms in the mind. He thus effectively replaced general learning principles with an internal grammar engine: a universal grammar wired into the human brain from birth. Almost like a physical organ, language would grow and develop, nurtured by the speaking community.
Chomsky Meets ChatGPT
Strikingly, much of Chomsky’s critique of Skinner could in principle apply to today’s large language models. But these systems follow Skinner’s recipe: they learn language without knowing what “life” is, or what it means to lose it. And yet, they produce surprisingly apt responses to threats, jokes, grief, or declarations of love. All based on patterns in language data.
This is behaviorism reimagined: language learned entirely through exposure, without intent or meaning. The model has no beliefs, no desires, no fears, and yet behaves as if it does. This is exactly what Chomsky thought impossible.
Chomsky’s focus, and the drive behind the school of Generative Grammar, was never about language in use, but about the internal conditions that make language possible, what he called competence. Still, in arguing against Skinner, he drew on examples from everyday understanding. And here the problem arises: if meaning depends on internal mental structure, how do we understand notes on the fridge, signs in windows, or books written long ago? We interpret them even when no mind is present behind the sign. Language works even when no one’s home.
Wittgenstein and the Language Game
While Chomsky looked inward toward innate structure, Wittgenstein turned outward. In Philosophical Investigations (1953), he introduced the idea of “language games” to show that meaning doesn’t reside in definitions or syntax, but in use. Words may stand for a lot of things, but meaning is what we do with them.
A language game is a social activity: greetings, promises, requests, commands, apologies, explanations and so on. These are practices with rules, which in these games are not fixed. They are learned through participation, not by being told the what and the how.
In this view, understanding is not so much a hidden mental state as it is something displayed through action. We don’t need to know what “your money or your life” means in the abstract; we only need to know how to respond.
This opens a door for AI. If a model acts appropriately in the language game, we may not need to ask whether it understands, instead we allow it to participate in our own language games.
AI as Behaviorist Machine — With Transparency
When people call language models parrots, they imply that these systems merely mimic, blindly regurgitating past patterns. We now know that they do more. They can make new sentences, and respond adequately to sentences that are newly formed.
Modern AI systems offer transparency: we can inspect attention weights, visualize neural activations, trace layer-by-layer representations. For a generative linguist from the 1980s, these patterns would look like echoes of deep structure. Ironically, when we finally build a purely behaviorist language model, it reveals structure everywhere; structures that resemble the syntactic trees from the 1960s.

Graph pulled out from a language model (NbAiLab/nb-bert-base) from its first layer for the sentence “John found a bike near the lake”. Connections based on attention weights. Note the hierarchical phrasal structuring. The labels on edges are added afterwards, substituting the attention weights.
And more than that, AI models act socially. They adapt tone, negotiate meaning, play roles, and we play along. They enter our language games through fluency without consciousness. Like a Wittgensteinian actor, they follow rules without knowing them.
Although we can tell the difference between meaning and its simulation (see our essay on AI and film), the distinction can be blurred.
The Metaphor Machine
One of the hardest problems in NLP has long been metaphor. From Lakoff and Johnson’s Metaphors We Live By to countless failed attempts to detect figurative language, researchers have struggled to capture what happens when we say “the idea caught fire” or “time is running out.”
And yet, language models shine here. They generate, interpret, and remix metaphors with astonishing ease, simply because they live in them. LLMs are trained in a metaphor-rich environment, and, for them, figurative language is the norm.
In a strange twist, the parrot turns poet. What once required human insight now emerges from statistical resonance. Importantly, the model doesn’t know; knowledge lies buried in the language it comes from, and the model has become language’s most sensitive echo chamber.
Parrot or Partner?
Weizenbaum’s warning echoes louder now. Modern LLMs are far more sophisticated than ELIZA, but our relationship remains: we assign meaning. We let them speak, advise, console, and narrate.
And still, amid all this, only one thing is certain:
AI may take our jobs, but it can’t take our vacations. Because to live a life is more than to model one.
메타데이터
- post_id
- 7b36f1aff44e
- slug
- knowledge-buried-in-language-7b36f1aff44e
- url
- https://medium.com/@yoonsen/knowledge-buried-in-language-7b36f1aff44e
- canonical_url
- https://medium.com/@yoonsen/knowledge-buried-in-language-7b36f1aff44e
- author_url
- https://medium.com/@yoonsen
- status
- ok
- fetched_at
- 2026-06-16 19:09:56