Notebooks 1 : Language Understanding
I decided to publish various notes I took while thinking about Alignment and Common Sense. I did not find a place for some of these…
Notebooks 1 : Language Understanding
I decided to publish various notes I took while thinking about my blog post, Alignment and Common Sense. I did not find a place for some of these thoughts, but I like many of them on their own. I have quite of few of these scattered thoughts (some I quite like, but many are rather poor quality). They are trying to extend that piece into the field of language understanding. I decided to just publish them in the tradition of Pascal or Wittgenstein, not that the thoughts are anywhere near as good.
In fact, I’m a bit embarrassed by how some passages mimic the late Wittgenstein. My stylistic imitation evinces a hubristic desire to emulate his brilliant way of thinking, which of course fails.
Despite my embarrassment, I greatly enjoyed this imitation as a writing practice. I have edited these notes a bit to keep out the really unintelligible blurbs. I will probably continue to publish other Notebooks passages, in case anything good comes up.
— — — — —
In the racing game, we know how to continue. In chess we do not know how to continue. The challenge is to decide what is the optimal path, precisely by thinking through various outcomes as an algorithm (running simulations). It should not be at all surprising that machines could surpass humans here — although for the Go community, I imagine it would feel like ChatGPT does for the general public.
So are language-games more like Go? Or like a video game? There’s a mix of them — some language-games require technical investigation, or an abstraction. Most depend upon some interpretation: a feeling for how one should continue in the present context.
Misalignment in the racing game just shows there is something beyond the explicit scoring function that we consider important to the game.
I remember seeing a viral football play where the quarterback calmly walked through the opponent linebackers, giving everyone a sense that the game was somehow paused, until he started running to the field goal. Was this play part of American Football? Was that a form of rewarding hacking? There are some potential exploits that had not been covered by any rules: those involving acting.
Should we be existentially concerned with misalignment because super-intelligence is a possibility (even as a remote one)? There are many normative constraints on our behavior that we barely notice but AI disregards them. For instance: do not lie about current events. AI will hallucinate or make stuff up: it is not breaking a grammatical, or even a logical, constraint but a normative one, ‘Do not lie.’ And there is a limit to our formalization of such problems — there is no complete formal definition of truth (Gödel). All the same, the limitation to a specific context cuts both ways. AI will not plan outside of its problem space. Many of the super-intelligence scenarios presuppose that AI develops some broader ontological understanding. If the AI can understand all the resources that go into a paper clip, including the social structures that underly this manufacturing, why would it be so difficult to explain basic legal constraints on behavior?
There is a contextual limit to the understanding of goals, which necessitates proxy metrics. This same limit would apply to intermediate goals.
Problems can be very misleading. The fact that language appears as a problem for modern computational linguists is still astounding. What do we mean by solving language?
Natural Language Understanding, as a single field, blurs together so many problems. It is like the tendency of scientists to see in their lab a microcosm for the entire world.
Understanding language was the holy grail because it would unlock thought — we think through language. Previously we though logical forms were behind language. These are no longer thought to be the essence of language, but they are not negligible. We expect speech to uphold certain logical norms in certain contexts. Logic is a constellation of our rules about language use.
Understanding language is to be able to use language. So is it using language? For a language model, language is not a means to anything — predicting speech is the game in itself. Asking if it uses language is like asking if we use the game of chess to play chess (sounds odd, but not necessarily false).
There was a time when offsides did not exist in games like Hockey. Kids playing in a frozen pond might have simply complained ‘hey stop sitting by our goal!’ The guilty party might see the opposing team try out the same strategy, and the game manifestly devolves into meaningless gimmicks.
There is something like a reflection on the general implications of our behavior. This reflection need not be consciously held — it could be born out in practice. Yet we see the meaning of a rule in the activity, as a generalizable experience (the shape of pieces is not part of the chess game — it is not something that can be generalized).
It was only later that the offsides rule would be formalized, as a painted line, followed by refs. They might even install camera or automate the tracking of a puck. The game still existed before these technologies; it was less efficient and cause more arguments. Hockey exists as an ideal game behind all these iterations, these different ways of putting it into practice (misleading picture).
Once hockey is fully standardized, we feel we could write it down and preserve it — there is a greater transmissibility in this formalization, an enframing of judgments. We even have video games with physics — not a complete physics, but everything that is relevant to recreating hockey.
This digitization of games is comparable to the development of AI. We cannot capture every rule of language use in a formal system (many language rules are spontaneous) so we make a new framework, like painted lines on the ice. We can then formalize those rules and guarantee them with the right techniques. In a way, we have captured meaning — a historian would be very blessed to encounter a usable image generator and see the visual associations of words. But does that mean the AI has arrived at some core of what words mean, has come closer to solving language? No more NHL 2K-something gets closer to ideal hockey with a more realistic physics engine. The reason for the game is still the experience of playing it. More details in the representation are irrelevant. Similarly, language is only grounded in the experience of using it.
A correlation between the brain and an AI model tells us little about either. You could also correlate a graph of the global economy with brain activity. Maybe we would do that to a clairvoyant.
Of course there are significant differences represented in the brain, which might correlate to signals in an AI model. That correlation only tells us that they are significant, not in what way. Neural activations are significant in an information theoretic sense.
Correlations only register events. Significance is purely subjective. In information theory, the events signify something else. The meaning of a message is not relevant to this framework.
메타데이터
- post_id
- 786d7ea4fc6a
- slug
- notebooks-1-language-understanding-786d7ea4fc6a
- url
- https://medium.com/@jadedjohn/notebooks-1-language-understanding-786d7ea4fc6a
- canonical_url
- https://medium.com/@jadedjohn/notebooks-1-language-understanding-786d7ea4fc6a
- author_url
- https://medium.com/@jadedjohn
- status
- ok
- fetched_at
- 2026-07-15 20:04:32