Anthropic Built Their Most Powerful AI and Then Said “Actually, You Can’t Have This”
We are so conditioned to watch AI companies release things first and explain themselves later. A model drops, the benchmarks come out, the…
Anthropic Built Their Most Powerful AI and Then Said “Actually, You Can’t Have This”
We are so conditioned to watch AI companies release things first and explain themselves later. A model drops, the benchmarks come out, the demos go viral, and the blog post about “responsible deployment” arrives six months after the fact when something has already gone wrong and everyone’s asking questions.
Claude Mythos was different.
In April, Anthropic released Mythos Preview through something called Project Glasswing, and they were pretty transparent about why the rollout was deliberately tiny. Internal evaluations had shown that Mythos-class models could find and exploit software vulnerabilities across every major operating system and web browser. Not because it was designed as a hacking tool. That’s the part that’s strange and worth sitting with. It wasn’t built for offensive security. It just became capable enough at reasoning that it could do it anyway.
Think about what that means for a moment. You build a model that’s really good at understanding code. Good enough that it can debug complex systems, refactor legacy codebases, reason about what code is doing. And somewhere on the capability curve, “understanding what code does” starts to include “understanding where code fails.” There’s no bright line between those two things.
Anthropic’s answer was not “we’ll add a filter.” Their answer was to create a whole separate program for releasing this thing responsibly.
Project Glasswing brought in Amazon Web Services, Apple, Google, Cisco, Microsoft, JPMorgan, and several other major infrastructure holders and gave them early access to Mythos Preview. The explicit goal was to let the defenders use the model first: to find and patch vulnerabilities before bad actors could use a Mythos-class model to find and exploit them. It’s a race condition, basically. You can’t un-build the capability, so you try to make sure the defense gets a head start.
There’s something almost poetic and slightly unhinged about this. A room full of the world’s largest tech companies sitting with an AI that is extremely good at finding holes in their own software, using it to close those holes before someone else opens them. The fox and the henhouse analogy doesn’t quite work because in this case the fox is helping design a better lock. I think.
U.S. intelligence agencies and government officials apparently started paying close attention after Mythos dropped. When a company voluntarily restricts access to its own product and the intelligence community is watching what happens next, you’re probably in interesting territory.
Now, two months later, Fable 5 is public. It’s built on the same underlying model as Mythos. The difference is that the specifically dangerous capabilities in the cybersecurity and biology domains have been muffled with additional safeguards, and Anthropic kept the unrestricted version, now called Mythos 5, inside Glasswing where it started.
They’ve essentially created a two-tier release strategy based on risk. The version with full capability stays with the people whose job is defense. The version with the specific dangerous parts turned down gets released to the rest of us. This is not how software companies usually operate. This is closer to how you’d think about responsible release of a genuinely dual-use technology.
I find the naming interesting too. The announcement noted that “Fable is from the Latin fabula, meaning that which is told, akin to the Greek mythos.” They’re related words. One is the sanitized version of the story that gets passed around. The other is the original.
Make of that what you will.
The honest take from my end: I used to be a little skeptical when AI companies talked about safety as a genuine priority, because the track record of those words matching actions is not great across the industry. What Anthropic did with Mythos is the closest thing I’ve seen to “we built something powerful, realized the implications, and actually paused instead of just shipping it.”
Whether that continues as the models get more capable is a different question. But the precedent of deliberately structuring access tiers around real-world risk, and being transparent about exactly why, is at least a more honest answer than most.
Also, somewhere in the world there is a small room of engineers at Cisco using an AI that isn’t available to the public to find bugs in the infrastructure the rest of us use every day. I find that very strange and also somewhat comforting, which is a sentence I did not expect to write.
메타데이터
- post_id
- d09d34d6e03f
- slug
- anthropic-built-their-most-powerful-ai-and-then-said-actually-you-cant-have-this-d09d34d6e03f
- url
- https://medium.com/@tusharkuradia/anthropic-built-their-most-powerful-ai-and-then-said-actually-you-cant-have-this-d09d34d6e03f
- canonical_url
- https://medium.com/@tusharkuradia/anthropic-built-their-most-powerful-ai-and-then-said-actually-you-cant-have-this-d09d34d6e03f
- author_url
- https://medium.com/@tusharkuradia
- status
- ok
- fetched_at
- 2026-06-27 18:20:27