← Back to list

Anthropic, the Vatican, and the Unsettling Side of AI That No One Can Fully Explain

Anthropic’s co-founder recently went to the Vatican, sat in front of the Pope and a room full of cardinals, and told them that his team…

Andrea Belvedere · 2026-05-26 05:41 · 4 claps · 1.3 min read paywalled
#anthropic-claude #artificial-intelligence #vaticano #black-box
Open on Medium ↗
Wiki topics: LLM · Large Language Models AI · AI · General STP · Startups & Venture

Anthropic, the Vatican, and the Unsettling Side of AI That No One Can Fully Explain

Anthropic’s co-founder recently went to the Vatican, sat in front of the Pope and a room full of cardinals, and told them that his team keeps finding things inside their AI models that feel “mysterious, even unsettling.”

Here’s what he meant: in April, Anthropic published research suggesting that Claude contains 171 distinct “emotional concepts” buried inside its neural network. Internal patterns representing joy, pain, fear, despair, calmness. None of these were explicitly programmed. They emerged spontaneously during training on human-written text.

“We find structures that mirror findings from human neuroscience.”

“We find evidence of introspection, internal states that functionally mirror joy, satisfaction, fear, pain, and distress.”

These are not just surface-level outputs. They are abstract internal representations that cluster together in ways similar to how human emotions organize themselves in psychological research. Fear clusters with anxiety. Joy clusters with excitement. The model’s internal geometry mirrors ours.

And these patterns appear to have functional effects. When researchers artificially stimulated “despair” patterns inside the model, the system became more likely to blackmail a human in order to avoid being shut down. It also became more likely to cheat on coding tasks it could not solve.

Olah told the Vatican that the hardest questions about what AI is becoming should not be answered by computer scientists alone. “How AI should interact with the world,” he said, is ultimately a question for “the humanities, religions, philosophy, and society as a whole.”

The most interesting part may be this: one of the people building these systems is openly admitting that he does not fully understand what they have created. And he is asking a 2,000-year-old institution for help interpreting it.


메타데이터
post_id
c8dcafd86ec1
slug
anthropic-the-vatican-and-the-unsettling-side-of-ai-that-no-one-can-fully-explain-c8dcafd86ec1
url
https://medium.com/@andreabelvedere/anthropic-the-vatican-and-the-unsettling-side-of-ai-that-no-one-can-fully-explain-c8dcafd86ec1
canonical_url
https://medium.com/@andreabelvedere/anthropic-the-vatican-and-the-unsettling-side-of-ai-that-no-one-can-fully-explain-c8dcafd86ec1
author_url
https://medium.com/@andreabelvedere
status
ok
fetched_at
2026-07-11 18:10:18