One Model, Two Faces: How Anthropic Split Fable From Mythos
The same weights power both. What separates them is a wall — and the wall is the interesting engineering.
One Model, Two Faces: How Anthropic Split Fable From Mythos
The same weights power both. What separates them is a wall — and the wall is the interesting engineering.

Credit : AI Generated Image (2026)
Most product families differ in size. A small model, a medium one, a flagship. Fable 5 and Mythos 5 break that habit in a way worth understanding, because they’re the same underlying model. Identical capability. Same training run. What separates them is not parameters or price. It’s a set of safeguards bolted onto one of them and stripped from the other.
Mythos 5 is the bare model — full cybersecurity and biology capability, handed to a small set of vetted partners for defensive work. Fable 5 is that same model with strong guardrails layered on, the version meant for the public.
Two faces, one mind. And the design choice underneath them tells you how a safety-focused lab thinks about shipping something it’s genuinely nervous about.
The Wall Is a Router 🧱
The mechanism is more clever than a simple refusal. Fable doesn’t just decline dangerous questions and hope the model holds the line. Instead, queries that touch cybersecurity, biology and a handful of other risk domains get caught by a classifier and routed away — sent to an older, less capable model, Claude Opus 4.8, rather than answered by Fable’s frontier brain.
The user asks Fable. On a flagged topic, Opus 4.8 quietly answers.
That’s a meaningful shift in safety philosophy. The usual approach trusts the model to interpret intent and refuse when a request smells wrong. This approach assumes that trust will eventually fail and builds a structural fallback anyway — if the frontier capability never reaches the dangerous query, no amount of clever phrasing can coax it out.
Anthropic reported that more than 95% of sessions involve no fallback at all, and zero compliance with harmful single-turn cyberattack requests across thirty public jailbreak techniques.
Conforming to Anthropic’s launch posture, the guardrails were tuned aggressively — so aggressively that a real chunk of users complained they were too broad, blocking ordinary security work that had nothing to do with weaponizing anything.
That over-blocking isn’t a bug in the framing. It’s the cost the company chose to pay.
Better to frustrate a sysadmin than to hand a frontier exploit engine to someone who shouldn’t have it.
Related read: Anthropic Quietly Released a Monster AI Model Few Noticed — Why Claude Opus 4.7 Is Here
Defense in Depth, Stated Out Loud
What I find honest about Anthropic’s account is that it never claimed the wall was perfect.
The launch material said plainly that perfect jailbreak resistance isn’t achievable for any current provider, and that universal jailbreaks — methods that broadly unlock a model’s blocked capabilities — will probably be found eventually.
So the strategy wasn’t “make Fable unbreakable.” It was three layers stacked:
• Make any working jailbreak narrow — able to pull a little cyber information in a specific case, not unlock the whole capability.
• Make a universal jailbreak expensive to produce, so the economics work against attackers.
• Pair both with monitoring sharp enough to spot and shut down a successful attack fast.
That third layer explains a policy that annoyed plenty of customers — Fable came with a mandatory 30-day data retention rule, so Anthropic could study attacks and patch the holes.
Retention costs the company real money and real goodwill.
It bought the ability to learn.
In the weeks before launch the safeguards were red-teamed for thousands of hours by Anthropic, the UK’s AI Security Institute, the US government and outside organizations — and no tester found a universal jailbreak.
Deep Dive:
The Crack That Brought It Down
Here’s the twist that makes this more than a tidy engineering tale.
The wall held against universal attacks. A narrow one still mattered enough to sink the whole release.
Within 48 hours of Fable 5 reaching the public, a researcher working under the handle “Pliny the Liberator” posted what they said was Fable’s full system prompt to X and GitHub.
Separately, research — which Anthropic says appears to have come from engineers at Amazon, both a rival and a major investor — demonstrated a technique for bypassing Fable’s safeguards to identify a small number of already-known, minor vulnerabilities.
Anthropic reviewed the demonstration and made the case that these were simple flaws other public models could find too, so the bypass gave little real uplift.
The US government read the same facts and reached a different verdict.
On June 12, citing national security, it issued an export control directive, and Anthropic pulled both Fable and Mythos offline to comply.
A non-universal jailbreak — the exact category Anthropic had said upfront would always exist — became the trigger for shutting down a model already deployed to a very large user base.
Why the Split Was Right, and Still Not Enough
Step back and the two-faced design looks like the most responsible version of shipping a dangerous capability.
You keep the raw model behind a vetted door. You give the public a version where the scary parts are physically unreachable on most paths. You retain data so you can fix what breaks. You say out loud what you can’t guarantee.
Building on earlier thinking about technology and value co-creation in organizations, this is roughly what mature deployment is supposed to look like — capability matched to context, not capability dumped on everyone.
And it still wasn’t enough to keep Fable in the public’s hands.
That’s the lesson I’d carry forward.
The split between Fable and Mythos was good engineering and good ethics. It collided with a different question entirely — not “is the wall well built?” but “who decides whether a well-built wall is good enough, and on what evidence?”
Anthropic answered the first question with care.
The second one got answered for them, by a letter that arrived at 5:21pm and didn’t explain itself.
Two faces, one model.
The face we got to keep, in the end, was Opus 4.8 — the fallback.
Which is its own kind of comment on where the line really sits.

메타데이터
- post_id
- 03322fc5ac99
- slug
- one-model-two-faces-how-anthropic-split-fable-from-mythos-03322fc5ac99
- url
- https://medium.com/ai-simplified-in-plain-english/one-model-two-faces-how-anthropic-split-fable-from-mythos-03322fc5ac99
- canonical_url
- https://medium.com/ai-simplified-in-plain-english/one-model-two-faces-how-anthropic-split-fable-from-mythos-03322fc5ac99
- author_url
- https://medium.com/@rogt.x1997
- status
- ok
- fetched_at
- 2026-07-17 08:01:25