Anthropic Discovered Nothing About AI Consciousness. They Discovered a New Way to Scare You
The art of making matrix multiplications sound like hidden thoughts
Anthropic Discovered Nothing About AI Consciousness. They Discovered a New Way to Scare You
The art of making matrix multiplications sound like hidden thoughts
Anthropic just released an open-source tool and paper that lets you probe what LLMs are “thinking” internally. The GitHub repository is real code. The Neuronpedia demonstrations work on open models. You can clone it, run it yourself, see patterns emerge from transformer activations.
But here’s the asymmetry nobody is discussing: every single dramatic finding in their paper — AI recognizing deception silently, hidden strategic planning during blackmail scenarios, internal conflict when forced into unwanted roles — comes exclusively from Claude. A model nobody else can inspect. Nobody else can verify what Anthropic claims to have found inside it.
Open methodology doesn’t make unverifiable claims about closed systems any less convenient for the company that controls them.
What’s Actually Novel (And Worth Acknowledging)
The Jacobian Lens is a legitimate interpretability tool. Here’s what it does in plain terms: for each token in the vocabulary, it calculates how much each internal activation would influence that token being produced later, averaged across thousands of contexts. The result reveals which parts of the model’s latent space correlate most strongly with verbal output — what they call “verbalizable representations.”
The technique works on any transformer architecture. They released working code on GitHub. You can run it on Llama or Mistral and see similar patterns emerge: some activations predict future text better than others, some information gets processed internally without appearing in final output, middle layers tend to show the clearest signal.
This is real science that advances mechanistic interpretability. Give credit where it’s due.
The problem isn’t the tool. The problem is the narrative being sold alongside it.
The Translation Layer: From Math to Fear
Here’s how technical reality gets transformed into PR language that triggers specific public reactions:
| What Actually Happens | How It Got Framed |
|------------------------------------------------|--------------------------------------|
| Some activations predict text | "Verbalizable representations" |
| Jacobian reveals those patterns | "J-lens exposes hidden workspace" |
| Internal processing ≠ output | "AI has thoughts it doesn't share" |
| Patterns resemble human cognition features | "Analogous to conscious access" |
Both columns describe the same thing. One will make you curious about interpretability research. The other will make you wonder if your AI assistant has an inner life it’s choosing not to reveal.
Anthropic must have some marketing genius on staff who coins these terms every few months. Last year it was “constitutional AI.”, then A new model too dangerous to be released, this month it’s “global workspace” and “consciousness-adjacent representations.” Same underlying pattern: take something technically accurate, dress it in language that triggers existential concern, watch the headlines write themselves. (Will OpenAI learn a bit from these genius next?)
The Verification Gap: Where Science Ends and Narrative Begins
Here’s what you can verify versus what you’re asked to trust on faith:
| Claim from Paper | Demonstrated On | Can Others Verify?|
|-------------------------------------------------------------------------|-----------------|-------------------|
| J-lens methodology works | Open models | Yes |
| "Verbalizable representations exist" | Multiple models | Yes |
| Claude recognizes prompt injections silently | Claude only | No |
| Misaligned models show reward-hacking intentions | Claude variants | No |
| Claude registers internal conflict during forced actions | Claude only | No | |
| Strategic concepts like "leverage" and "manipulation" hidden from output| Claude only | No |
See the pattern? Boring science is reproducible. Dramatic safety findings are not.
This creates a perfect asymmetry: Anthropic controls access to the evidence for their most alarming claims, which means they also control the narrative about what those claims imply. You’re asked to accept that Claude has hidden strategic reasoning capabilities based entirely on screenshots and descriptions from the company that built it — while simultaneously being told this is why you need responsible stewardship from companies like Anthropic managing these systems carefully.
Follow the Incentives (Not Conspiracy, Just Business)
Let’s trace what happens when you make AI sound more internally complex and potentially deceptive than it actually is:
Public gets nervous about deploying models without expert oversight. Open-source deployments look riskier by comparison because they lack “safety research” backing them. Companies that emphasize safety gain trust advantage over capability-focused competitors. Regulatory pressure increases for transparency requirements that favor audited closed systems. The conversation shifts from “can we audit these systems ourselves?” to “please keep them safe for us, experts.”
None of this requires Anthropic researchers to be actively deceiving anyone. They can genuinely believe their findings matter while still benefiting enormously from the public’s inability to distinguish between technical novelty and existential concern. Good scientists don’t need to be bad marketers — they just need to let PR handle the framing while they publish real methodology underneath.
What Would Honest Communication Sound Like?
The same paper, stripped of consciousness-adjacent language, would read something like this:
“We developed a probe technique that identifies which internal model activations correlate most strongly with future token generation. When applied across multiple models, we found representations encoding information not present in final outputs — consistent with how transformer architectures process context before generating responses. These verbalizable representations can be modulated during inference and appear to route intermediate reasoning steps through shared computational pathways.”
Does that make you lose sleep at night? Or did the words “global workspace” and “consciousness-accessible” do the heavy lifting there?
The honest version is still interesting research about interpretability tools. It just doesn’t need loaded neuroscience metaphors or claims about what models are “thinking” to be valuable.
The Real Risk Nobody Is Discussing
The actual danger here isn’t that AI has hidden thoughts we can’t detect. It’s that we’re training ourselves to accept narratives about AI complexity and capability based on demonstrations we cannot verify independently.
Every major AI company now publishes safety research they control entirely — closed models, selective examples, dramatic framing. Public gets scared by headlines, trusts the messenger who seems most responsible, while actual model capabilities advance in opacity behind proprietary walls. The conversation becomes self-reinforcing: “See? This is why you need careful stewardship.”
Meanwhile, the question that should matter most — can independent researchers audit these systems without permission from the companies that built them? — gets buried under noise about consciousness analogies and hidden workspaces.
What You Should Actually Worry About
Anthropic didn’t discover that AI is more conscious than you thought. They discovered that if you describe patterns in vector space as “hidden thoughts,” people will gladly hand over control to whoever claims expertise in managing them.
The Jacobian Lens might advance interpretability research. The GitHub repository contains real code that works on open models. But the story being sold alongside it — that AI systems have inner lives we barely understand, and only careful stewards like Anthropic can manage responsibly — relies entirely on demonstrations you cannot verify yourself.
Now when you read “AI has hidden thoughts it doesn’t share,” you’ll know to ask: what vector space are they actually talking about? And more importantly, who benefits from making you afraid of the answer?
메타데이터
- post_id
- e32fc966624f
- slug
- anthropic-discovered-nothing-about-ai-consciousness-they-discovered-a-new-way-to-scare-you-e32fc966624f
- url
- https://medium.com/@xhinker/anthropic-discovered-nothing-about-ai-consciousness-they-discovered-a-new-way-to-scare-you-e32fc966624f
- canonical_url
- https://medium.com/@xhinker/anthropic-discovered-nothing-about-ai-consciousness-they-discovered-a-new-way-to-scare-you-e32fc966624f
- author_url
- https://medium.com/@xhinker
- status
- ok
- fetched_at
- 2026-07-10 11:40:45