8 min read

Anthropic expected six months of lead on Claude Mythos. It got twenty days.

A glasswing butterfly resting on a green leaf, its transparent wings revealing the leaf surface beneath
The glasswing butterfly is named for wings that show exactly what lies behind them. Anthropic chose the name deliberately. The question is what Glasswing the project is designed to reveal, and what sits outside its field of view. - Photo by Ben Berwers / Unsplash
Last updated: 21 August 2026. What's changed: the export controls were lifted on 30 June and Fable 5 returned globally on 1 July. An earlier version of this article said Fable 5 remained frozen, which stopped being true the day after it was written.

The US export controls on Claude Fable 5 and Claude Mythos 5 were lifted on 30 June 2026, nineteen days after they landed. Fable 5 returned to general availability worldwide on 1 July. Mythos 5 went back to a set of approved US organisations and stayed there.

The order had followed a report from Amazon researchers describing a way to prompt Fable 5 past one of its cybersecurity safeguards, in one case producing code that demonstrated how a vulnerability could be exploited. Anthropic reviewed the technique with Amazon and the government and concluded it exposed no unique Mythos-level cyber capability, calling it a borderline case for Fable 5's safeguards. It retrained the classifier, which now blocks the technique in more than 99% of cases, and said the stronger protection would increase false positives on legitimate coding work.

Katie Moussouris, CEO of Luta Security and a former adviser to the US government on export controls, reviewed the report Anthropic shared. She said the technique had been tested on known CVEs and on code with deliberately inserted vulnerabilities, and that no new vulnerabilities in real code appeared in the paper.

Her reading was that the response misread the paper. Unintended regulatory consequences arrive, she argued, when too few experts are advising on how closely defence resembles offence in cybersecurity.

The April disclosure had been taken seriously enough that Powell and Bessent convened the CEOs of America's largest banks at short notice, JPMorgan among the Glasswing launch partners.

On 2 June, three weeks before the freeze, Anthropic published a forecast of its own. Announcing the expansion of Project Glasswing to roughly 200 organisations, it said it expected rival developers to field models with comparable cyber capabilities within six to twelve months, potentially released without equivalent safeguards against misuse.

OpenAI released the full GPT-5.5-Cyber twenty days after that forecast, on 22 June, while both Anthropic models were still switched off. Google DeepMind followed with Gemini 3.5 Flash Cyber on 21 July, at forty-nine days.

The export controls spent nineteen days guarding a lead that had already gone.


What a defender can actually get

Claude Mythos 5 is not available to you. Project Glasswing runs at roughly 200 organisations, and Anthropic's stated criterion for admission is that a successful attack on the partner's codebase could affect more than 100 million people.

What is available is a set of application-based programmes that switch off refusal behaviour on general-purpose models. Anthropic runs the Cyber Verification Program for Opus and Sonnet class models. It is free, the review decision target is two business days, and approval covers the dual-use category only: prohibited activities such as ransomware development and mass data exfiltration stay blocked whatever your status.

Three rules decide whether it reaches you. Approval attaches to a specific organisation ID, so there is no route for an individual consultant and no carry-over from a team organisation to a personal workspace. Organisations on Zero Data Retention are not eligible, and the programme does not run on Google Vertex AI at all.

OpenAI splits the same idea across two tiers and admits individuals as well as organisations. Daybreak Blue removes the production screening layer from GPT-5.6 Sol. Daybreak Red gives approved users GPT-5.6-Cyber, trained specifically to refuse less on exploit-chain development, authentication bypass and privilege escalation.

OpenAI published its own measurement of the gap between them on 10 August. On an internal benchmark covering those three task types, GPT-5.6 Sol completes 1.5% of requests. With Daybreak Blue access it completes 2.0%. GPT-5.6-Cyber, behind Daybreak Red, completes 95.0%.

Daybreak Blue is the tier OpenAI recommends as the starting point for most defenders. It still refuses that work almost every time.

Getting past those refusals wants organisational vetting, a hardware security key on every individual account from 1 September, and enrolment that now also runs through Amazon Bedrock for teams already building there.

Google DeepMind put Gemini 3.5 Flash Cyber behind the narrowest gate of the three. It runs as a limited-access pilot for governments and trusted partners, reached through CodeMender, with expansion promised and no date attached.

← Scroll to see full table

Programme Who can apply Where it runs What it unlocks
Daybreak Blue (OpenAI) Individuals and organisations OpenAI platform, Amazon Bedrock Screening layer removed from GPT-5.6 Sol
Cyber Verification Program (Anthropic) Organisations only, admin-scoped. ZDR excluded Anthropic first-party, Bedrock (not Opus 5), not Vertex Dual-use unblocked on Opus and Sonnet
Daybreak Red (OpenAI) Individuals and organisations, stricter vetting OpenAI platform, Amazon Bedrock GPT-5.6-Cyber
Project Glasswing (Anthropic) Invitation, roughly 200 organisations Bedrock gated preview and first-party Claude Mythos 5
Gemini 3.5 Flash Cyber (Google) Governments and trusted partners CodeMender Limited pilot
Open weights No application Your hardware Whatever you can run

Sources: Anthropic Cyber Verification Program documentation; OpenAI Daybreak access tiers, 10 August 2026; Google DeepMind Gemini 3.5 Flash Cyber announcement, 21 July 2026.

The data-handling rule turned into an argument between the two vendors on 19 August. Anthropic runs a 30-day retention policy for business customers on covered models and says retention is necessary for security. OpenAI said the same day that it believes it can serve frontier models to businesses without retaining their data, previewing Private Safety Processing to spot misuse across related interactions while keeping content on customer infrastructure or encrypted under customer-held keys.

OpenAI's Aleah Houze gave the worked example: a user asking about a weakness in one company's software in one conversation, then asking about remote access and detection tooling in another. Cross-session cyber reconnaissance is the exact pattern, and it is also indistinguishable from a defender working a finding across two sessions.

For a security team in health or financial services, that argument decides which lane exists. Under Anthropic's terms, an organisation that cannot accept retention is outside the CVP entirely.

Underneath all of it sits a lane that asks nobody's permission. Hugging Face ended up there mid-incident, running GLM 5.2 self-hosted after the commercial models declined to read its attacker's payloads. Daniel Fox Franke, tracking a segmentation fault in ripgrep, was blocked by GPT-5.6 Sol's cybersecurity classifier and finished the work on GLM 5.2 and Kimi K3.

Both are Chinese open weights, from Z.ai and Moonshot. The export controls were written partly to keep frontier cyber capability away from China. The defenders blocked by American safety classifiers went to Chinese models to get their work done.

The Bengio-led International AI Safety Report puts open-weight capability less than a year behind the closed frontier. That gap is the price of that lane, and it may compress the way Anthropic's lead over rival labs did.


The two layers no programme covers

Not everyone read the April disclosure as a watershed. Marcus Hutchins, who stopped WannaCry and is now principal threat researcher at Expel, made the substantive version of the objection: attackers have long relied on social engineering and phishing to get in without ever needing a novel vulnerability.

Most breaches do not start with a zero-day. They start with a phishing email, a misconfigured service account, an MFA prompt approved at the wrong moment, or a helpdesk agent talked into a password reset.

AI accelerates most of those. Spear phishing that used to need skilled manual targeting can be produced at volume with accurate contextual detail pulled from automated OSINT, and automated reconnaissance maps misconfigured assets faster than any human red team.

The CSA and SANS Mythos-ready briefing, co-authored by 250 CISOs with contributors from Google, NSA and CISA, reached the same conclusion from the data. Time-to-exploit has collapsed to under a day in 2026 on Sergej Epp's Zero Day Clock. The impact of exploitation has not risen to match.

The most consequential incidents of recent years ran on credential abuse, social engineering and supply chain compromise rather than on zero-days. The briefing frames the clock as a leading indicator of where attacker capability is heading, not a measure of current damage.

The second layer is expanding while the argument stays on the first.

Antiy CERT documented 1,184 malicious skills across ClawHub before coordinated disclosure. Trend Micro found 492 MCP servers exposed to the internet with no authentication controls. Check Point Research disclosed remote code execution in Claude Code through poisoned repository configuration files.

OpenClaw: When AI Agent Marketplaces Become a Supply Chain Risk

Public code pushes to GitHub grew 78.4 per cent in the year to March 2026, against 6.6 per cent two years earlier, and the first quarter of 2026 alone brought 15.7 million new developer accounts. Both figures are floors, since the dataset covers public repositories only and GitHub excludes accounts pushing at volumes no person could produce.

Nothing in that data says what security tooling any of those accounts run. Broken Access Control has meanwhile overtaken Injection as the most common CodeQL alert, which is the defect shape you get when generated code skips the authorisation check.

The tooling that catches that would have to run by default from the first week an account exists. The most capable version of it now sits behind an application form.


Defending all three layers at once

The response has to cover the discovery problem the access programmes were built for, the agentic estate already producing incidents, and the human and configuration layer that stays the dominant breach path whatever a model can now do with a zero-day.

Step back and they resolve into one surface. Attacks chain across discovery, identity and the agentic estate, which is why single-layer defence does not hold. Phishing-resistant MFA, access review, dependency hygiene and agent governance is where real attacks live.

This is exposure management applied to a new moment. The vulnerability-centric reflex is a hangover from the Conficker and WannaCry era, when the unpatched CVE was the breach. The threat moved to identity, configuration and the agentic estate, and the lens has to move with it.

AI review belongs in that list with the same caution. Testing three AI security reviewers against one codebase showed each had different blind spots, and what gets caught depends on what you point the tool at. The judgement on whether a finding is real, and whether the fix closes the exposure, stays human.

The export controls lasted nineteen days. What replaced them is an application form, an organisation ID, a data-retention clause and a hardware key, administered by three companies rather than one government.

Access is the part of this that resolves itself. The lead Anthropic estimated in months lasted weeks, and expect open weights sit inside a year of the frontier.

Google spent April arguing that a strong general model beat a specialised one, then shipped Gemini 3.5 Flash Cyber in July, built on Flash rather than a frontier model precisely because the large cyber models are costly to run.

Vetting programmes are what an industry builds around a capability while it is scarce and expensive. Both conditions are easing, even as the programmes themselves tighten.

Neither version of the gate has ever governed the phishing email or the exposed MCP server. A model will help you with both. Checking whether your own server requires authentication has never needed an application form.


References

CyberDesserts Research