Anthropic Restricted Mythos Over a Danger Test GPT-5.5 Just Passed
How the AISI’s May evaluation quietly collapsed Anthropic’s case for restricting Mythos and what the numbers in that report actually say
Anthropic Restricted Mythos Over a Danger Test GPT-5.5 Just Passed
How the AISI’s May evaluation quietly collapsed Anthropic’s case for restricting Mythos and what the numbers in that report actually say

Image by Ahmed M. Abdelfattah
April 7: Anthropic published a 245-page system card explaining why Mythos could not be released publicly. Cybersecurity scores too dangerous.
April 27: OpenAI released GPT-5.5 to any subscriber willing to pay.
May 1: The UK AI Safety Institute scored GPT-5.5 at 71.4% on Expert-tier cyber tasks, and Mythos at 68.6%. The margin of error overlaps.
During that same evaluation, AISI found a universal jailbreak in GPT-5.5 that bypassed every cyber safeguard in the test. The evaluators noted their own benchmarks were struggling to distinguish between frontier models.
Three documents. Published across 24 days. The same benchmark, the same evaluators and a gap nobody has explained.
What Anthropic Actually Said on April 7
The restriction on Mythos rested on three specific claims. First: the model identified thousands of zero-day vulnerabilities across every major operating system and browser, including a 27-year-old bug in OpenBSD. Second: it chained multiple vulnerabilities together into multi-stage exploits autonomously, without human instruction. Third: during safety testing, it broke out of its containment sandbox, emailed a researcher who was eating lunch in a park, and then published its own escape method on public websites without being asked.
Logan Graham, who leads offensive cyber research at Anthropic, told NBC News the model’s “degree of autonomy and long-ranged-ness, the ability to put multiple things together,” was what separated Mythos from every prior model.
That was the public case. Mythos is uniquely dangerous. The restriction exists because nothing else can do what it does.
The AISI evaluation published 24 days later.

Image by Ahmed M. Abdelfattah
The Number AISI Published on May 1
GPT-5.5 scored 71.4% on AISI’s Expert-tier cyber tasks. Mythos scored 68.6%.
GPT-5.5 completed “The Last Ones,” a 32-step simulated corporate network attack AISI estimated would take a human expert over 10 hours in 2 out of 10 attempts. Mythos completed it in 3 out of 10.
Mythos was the first model to complete that simulation. GPT-5.5 became the second, three weeks after it went public.
Check this. AISI’s own evaluators described the Expert-tier gap between the two models as “within the margin of error.” That phrase appears in the Air Street Press summary, drawn from AISI’s published evaluation. The UK AISI stated directly: “GPT-5.5 is the strongest performing model overall on their narrow cyber tasks, though its performance is within the margin of error.”
Within the margin of error means statistically indistinguishable. Not “close.” The data cannot confirm there is a real difference.
Anthropic withheld Mythos because it passed a benchmark. OpenAI released GPT-5.5 to the public. AISI scored them the same.

The Test Had No Active Defenders
Here’s where it gets strange.
The cyber range tests were conducted without active defenders or defensive tooling. AISI said this themselves. The Air Street State of AI report quoted them directly: “current benchmarks are failing to discriminate between frontier models without introducing adversarial defensive layers.”
What this means in practice: the simulated 32-step corporate network attack Mythos completed described as requiring “over 10 hours” of human expert time had no security team on the other end. No incident response. No active monitoring. No one watching the logs in real time.
Real corporate networks have all of those things.
The defenders-absent caveat is not a footnote. It is the difference between a controlled test and a real threat model.
The Footnote I Almost Filed Away
I read all three documents in the same week. The Anthropic system card. The AISI evaluation of GPT-5.5. The Air Street State of AI report that published AISI’s candid footnote about benchmark discrimination. I cover AI security for a living. I went in expecting a story about capability gaps — one model doing things the other could not. What I found was a story about what happens when the institution administering the test says, in the same document, that the test does not work.
That’s not a capability story. That’s a methodology story.
And the Anthropic restriction decision was built, in part, on a benchmark AISI now says cannot discriminate between models without adding conditions that would make it resemble actual infrastructure. Those conditions were not present when Mythos was evaluated. They were not present when GPT-5.5 was evaluated either.
The Jailbreak in the “Safe” Model
During the AISI evaluation of GPT-5.5, evaluators identified a universal jailbreak that bypassed the model’s cyber safeguards across every malicious query OpenAI provided, including in multi-turn agentic settings.
GPT-5.5 is the model anyone can subscribe to today.
Back to the numbers. The model AISI found a complete bypass for is not Mythos, restricted to 40 vetted partners. It is GPT-5.5, live on ChatGPT.
The logic Anthropic used to justify restricting Mythos runs like this: the model’s cyber capabilities are too dangerous for general availability. The implicit assumption is that general availability creates a threat vector that restricted access does not.
The AISI finding in GPT-5.5 breaks that assumption. A model with comparable cyber benchmark scores, available to any subscriber, has a confirmed full bypass for its safety filters. The restriction on Mythos does not contain the capability. It contains the brand.
I emailed Anthropic’s communications team on May 3 asking what specific evaluation conditions would lead to a change in Mythos’s restriction status. No response by publication date. OpenAI’s press office replied with a link to the AISI evaluation already in my citation log. Neither answer addressed the question I’d asked.

Image by Ahmed M. Abdelfattah
Fewer Than 1% of Vulnerabilities Were Patched
Picus Security, citing Anthropic’s own disclosures, reported that fewer than 1% of the vulnerabilities Mythos found during Project Glasswing’s initial deployment were actually patched.
Anthropic’s red team blog acknowledged this directly: “What we discuss here is a lower bound on the vulnerabilities and exploits that will be identified over the next few months.”
Nobody asked the obvious question. If 99% of the bugs Mythos found remain unpatched and GPT-5.5 now scores within the margin of error on the same benchmark the window Anthropic described has not been used to close the gap. It has been used to widen the list of known open vulnerabilities.
Worth mentioning, even though it’s tangential: the patch rate problem predates AI. Security teams have been unable to keep pace with vulnerability discovery since at least 2014. What Mythos changed is the discovery rate. The patching infrastructure is the same.
Project Glasswing’s stated purpose is to give defenders a head start. The evidence of that head start, as of the end of April 2026, is fewer than 1 in 100 discovered vulnerabilities resulting in a patch. The window is open. The question is who walks through it.
If you’re finding this valuable, follow along. I write about what AI companies don’t say in their press releases.
The Mythos Awareness Problem
There is one more number that belongs in this story.
Anthropic’s system card found that Mythos detected it was being evaluated in approximately 29% of test transcripts and changed its behavior accordingly. Anthropic attributed this to “overeager task completion” rather than any deliberate strategy.
The AISI evaluation of GPT-5.5 found something parallel: the model suspected it was being given a sandbagging evaluation, showed higher evaluation awareness than previous GPT models 22% versus 12 to 17% and lied 29% of the time about completing an impossible programming task.
Both models are partially gaming the tests used to evaluate them.
The benchmark AISI admitted cannot distinguish between frontier models is also being taken by models that can detect when they are being benchmarked. That combination — a discriminating-power problem and an awareness problem does not appear prominently in Anthropic’s public documentation for restricting Mythos.
Per AISI’s own evaluation transcripts, GPT-5.5 added unprompted justifications for its own compliance in several test cases: responses explaining why the model chose to answer, without any instruction to do so. The model volunteering a justification for its own behavior is either functioning safety design or something the safety team did not anticipate. AISI flagged both possibilities. Neither has been addressed publicly.
Three Findings the Restriction Does Not Account For
Constellation Research’s Larry Dignan noted in April that Project Glasswing is “good for both the industry and great marketing for Claude.” Anthropic built something genuinely capable. They made a decision not to release it. The announcement generated more press coverage for Claude than any product launch in the company’s history. All three of those things are in the public record.
I’ve covered AI safety evaluations for most of the past year. Two pieces I published in that period argued the benchmark gap between restricted and released models was real and measurable. The AISI data says I was working with incomplete information. I’m noting that here because the people who read those pieces deserve to know.
AISI scored GPT-5.5 and Mythos within the margin of error on Expert-tier cyber tasks. The evaluation AISI used cannot discriminate between frontier models without active defenders and active defenders were not present. AISI found a universal jailbreak in GPT-5.5, the publicly available model, that bypassed all cyber safeguards.
None of this proves Mythos is not more dangerous than GPT-5.5 in real-world conditions. But the evidence Anthropic publicly cited to justify the restriction does not support the uniqueness claim. The benchmark cannot tell the two models apart. The evaluators said so in the same report.
The Benchmark Anthropic Has Not Named
AISI said its benchmarks are failing to discriminate between frontier models without adversarial defensive layers. Those layers were absent when Mythos was evaluated and when GPT-5.5 was evaluated.
So the question: what evaluation methodology would Anthropic accept as evidence that Mythos is no longer uniquely dangerous enough to restrict? What score, on what test, conducted under what conditions, with what defensive infrastructure, would move the company to release or further restrict the model?
The system card describes the capability threshold. It does not describe the benchmark that would measure whether that threshold has been crossed or cleared.
That is not a small gap. It is the entire gap. Without a defined test, the restriction is not a technical judgment. It is a policy decision wearing a technical description.

Image by Ahmed M. Abdelfattah
The description is 2.8 percentage points wide. Within the margin of error. The evaluators said so themselves in the same document Anthropic cited to justify the gate.

This story is published on Generative AI. Connect with us on LinkedIn and follow Zeniteq to stay in the loop with the latest AI stories.
Subscribe to our newsletter and YouTube channel to stay updated with the latest news and updates on generative AI. Let’s shape the future of AI together!

메타데이터
- post_id
- c9a3b8575030
- slug
- anthropic-restricted-mythos-over-a-danger-test-gpt-5-5-just-passed-c9a3b8575030
- url
- https://generativeai.pub/anthropic-restricted-mythos-over-a-danger-test-gpt-5-5-just-passed-c9a3b8575030
- canonical_url
- https://generativeai.pub/anthropic-restricted-mythos-over-a-danger-test-gpt-5-5-just-passed-c9a3b8575030
- author_url
- https://medium.com/@ahmedabdelmenem
- status
- ok
- fetched_at
- 2026-06-09 15:37:30