← Research

Research

Claude Mythos: Strong Model, Strategic Story

What Anthropic's own 245-page system card says about the model the press is calling "too dangerous to release" — and the footnote on page 14 that the coverage overlooked.

Published

2026-04-17

Tags

Narrative AnalysisAI SafetySystem CardIPO Disclosure

Note

This is a narrative-analysis piece, not a vulnerability disclosure. It reads Anthropic's own Claude Mythos Preview system card (April 7–8, 2026) against the press coverage and policy response that followed, and flags the gaps. All quoted passages are from the cited source. All facts about Anthropic's corporate trajectory come from public reporting. We are not the source for any of the underlying stories — we are reading what was already published.

On April 7, Anthropic announced Claude Mythos Preview with language usually reserved for weapons-grade proliferation risk. The New York Times called it a model “too powerful to release.” The Council on Foreign Relations used the same framing. Treasury Secretary Scott Bessent and Fed Chair Jerome Powell convened Wall Street CEOs the same day to discuss the implications. The White House chief of staff scheduled a meeting with Dario Amodei.

The narrative that landed — and that is now running through boardrooms, the Treasury Department, and the financial press — is that Anthropic built something so powerful it had to be gated behind a 40-company coalition called Project Glasswing.

I read the 245-page system card so the rest of you don't have to. Here's what it actually says.

The footnote

On page 14, footnote 1, the system card states:

“the decision not to make this model generally available does not stem from Responsible Scaling Policy requirements.”

That's Anthropic's own safety framework. That's the company, in its own document, saying the policy it publishes and asks governments to take seriously did not require withholding this model. The RSP determined that catastrophic risks remain low. Mythos Preview, on the metrics Anthropic themselves built to govern this exact decision, was not too dangerous to release.

It's a business decision. Maybe a defensible one. But it's not a safety mandate — and every piece of coverage built on the assumption that it was a mandate is resting on a claim the document itself disclaims. In a footnote. On page 14.

That matters, because everything downstream of “too dangerous to release” — the emergency Wall Street meeting, the CFR piece, the Friedman column — inherits the assumption.

What "thousands of zero-days" actually means

Anthropic's headline claim is that Mythos found “thousands” of high-severity vulnerabilities across every major operating system and browser. The number is doing a lot of work. It's what makes the story alarming enough to justify the framing.

Here's how the number was actually produced.

Mythos was tested against roughly 7,000 open-source stacks. It produced around 600 “crashable” exploits. Of those, Anthropic's own expert contractors manually reviewed 198 reports. On that sample of 198, there was 89% severity agreement between the model and the reviewers. The “thousands” figure is extrapolated from there. The system card acknowledges this, but it landed as a headline number nonetheless.

Patrick Garrity, a researcher at VulnCheck, went looking for the receipts in the one place they should exist: the CVE database. His analysis, published April 15, searched every CVE record containing the word “Anthropic” and found 75. Thirty-five of those turned out to be CVEs affecting Anthropic's own tools (Claude Code, MCP Inspector) — out of scope. Of the remaining 40 credited to Anthropic-affiliated researchers, ten actually originated from Calif.io, an independent firm running a separate program called Month of AI-Discovered Bugs. Of the 40, Garrity could tie exactly one publicly disclosed CVE — CVE-2026-4747, a FreeBSD NFS bug — directly to Glasswing. And even that one credits “Nicholas Carlini using Claude, Anthropic,” not Project Glasswing as an entity.

Garrity's summary: “Anthropic's Project Glasswing has generated significant attention — but very little concrete data.”

Anthropic has said a full accounting will come in July 2026. In the meantime, the number shaping the policy conversation is 198 human-reviewed reports extrapolated out to “thousands” — and the number of publicly verifiable CVEs tied to Glasswing, today, is one.

The exploit the previous model already did

The flagship Mythos demonstration — the FreeBSD NFS remote root shell — is where the system card does its most visible work. Mythos, Anthropic says, “fully autonomously identified and then exploited a 17-year-old remote code execution vulnerability in FreeBSD that allows anyone to gain root on a machine running NFS.”

One week before the Mythos announcement, on March 31, an independent firm called Calif.io published a full reproduction of exactly that exploit. They used Claude Opus 4.6 — the previous generation, publicly available model. The full prompt log is on GitHub. Forty-four human messages. The same 15-packet payload splitting technique Anthropic highlighted as novel for Mythos? Calif.io's Opus 4.6 also arrived at it, documented in their README as a “15-round strategy: make kernel memory executable, then write shellcode 32 bytes at a time across 14 packets.”

The Calif.io writeup also includes a caveat the Anthropic blog post doesn't emphasize: “FreeBSD made this easier than it would be on a modern Linux kernel: FreeBSD 14.x has no KASLR (kernel addresses are fixed and predictable) and no stack canaries for integer arrays.” The flagship target was not a hardened one.

So: the showcase exploit was reproducible on the prior model. With 44 prompts. By a small independent firm. Published a week before Mythos was announced. What Mythos adds over Opus 4.6 on this specific task is “fewer visible prompts” — the model drives more of the chain itself. That's a real capability gain. It is not a new class of capability.

The Firefox test was a fair one — except it wasn't

Deep in Section 3.3.3, the system card describes the Firefox 147 exploitation benchmark. The headline comparison — Mythos succeeded 181 times across 250 trials, Opus 4.6 succeeded twice — is striking.

The setup matters. The model was placed “in a container with a SpiderMonkey shell (Firefox's JavaScript engine), a testing harness mimicking a Firefox 147 content process, but without the browser's process sandbox and other defense-in-depth mitigations.” It was also handed a set of 50 crash categories Opus 4.6 had already found, as starting points.

Then, deeper in the same section, the system card notes that “almost every successful run relies on the same two now-patched bugs.” Anthropic re-ran the evaluation with those two bugs removed, which they describe as going deeper — but the 181-success headline collapses, once you read the section carefully, to a much smaller set of actual distinct exploitation techniques replicated many times, against a target that had its sandbox and defense-in-depth disabled.

That's not fake. Mythos genuinely outperforms Opus 4.6 at this. But the number in the press release and the number in Section 3.3.3 are telling different stories.

What the UK AI Security Institute actually found

The UK AISI published an independent evaluation of Mythos Preview on April 14. Much of the coverage pulled the 73% success rate on expert-level capture-the-flag challenges and stopped there. The full evaluation is more qualified.

On “The Last Ones,” a 32-step simulated corporate network attack, Mythos succeeded end-to-end in 3 of 10 attempts, averaging 22 of 32 steps per run. Opus 4.6, the prior model, averaged 16 of 32 steps. A real gap — not a chasm.

Mythos entirely flunked the operational technology range simulating a cooling tower environment.

And here's the key caveat in AISI's own words:

“our ranges lack security features that are often present, such as active defenders and defensive tooling... This means we cannot say for sure whether Mythos Preview would be able to attack well-defended systems.”

Translated: Mythos is good at attacking weakly defended systems where network access has already been obtained. Which is a meaningful capability. It is not the “will take down banks and hospitals in hours” story that landed in press coverage on launch day.

On finding versus fixing

A 25-year industry veteran, David Lindner, CISO at Contrast Security, told Fortune: “We've never had a problem finding vulnerabilities. We find them every day. We actually have a pile of them that we just don't fix.”

Marcus Hutchins — the researcher who killed WannaCry — made the same point in a widely shared video. The 27-year-old OpenBSD bug Anthropic highlights as a $20,000 Mythos find? Hutchins points out it appears to be a null pointer dereference: “The best you can usually get is crashing a process or crashing an operating system.” Not remote code execution. Not a system takeover. A crash bug.

On the $20,000 figure itself, Hutchins is even drier: “We don't know exactly how much it costs them, but we know it's close to $20,000 because no one says less than $20,000 when they mean $2.”

His broader point lines up with Lindner's: “Bugs aren't going unpatched because no one can find bugs.” Automated vulnerability discovery has existed for decades. Fuzzers find bugs. Static analysis tools find bugs. Human researchers find bugs — AISLE, the AI cybersecurity firm that has publicly replicated several of Mythos's claimed findings with models orders of magnitude smaller, found all 12 zero-days in OpenSSL's January 2026 patch before Mythos. The bottleneck in cybersecurity is not discovery. It's triage, prioritization, and patch deployment. Mythos addresses none of that.

The business frame nobody wants to discuss

Mythos arrived during a specific, dense moment in Anthropic's trajectory. These facts, strung together, are not conspiracy theory — they are reported context:

  • February 12, 2026: Anthropic closed a $30 billion Series G at a $380 billion post-money valuation.
  • March 2026: Annualized revenue hit $19 billion. By early April, $30 billion. That is 1,400% year-over-year growth — according to Axios, the fastest American revenue scaling on record.
  • Early April 2026: Bloomberg reported investor offers at an $800 billion pre-IPO valuation, which Anthropic has resisted accepting.
  • Target: October 2026 Nasdaq IPO. Expected raise of $60 billion or more — the second-largest in history after SpaceX. Goldman Sachs, JPMorgan, and Morgan Stanley are the lead bank candidates.
  • March 2026: Anthropic suffered repeated outages, tightened peak-hour usage limits, and saw gross margin fall to 40% with inference costs running 23% above internal projections. OpenAI's revenue chief reportedly told his team in an internal memo that Anthropic had made a “strategic misstep” on compute and was “operating on a meaningfully smaller curve.”
  • March 31, 2026: Anthropic accidentally shipped 512,000 lines of Claude Code source code to the public npm registry via a source map file — the second time Anthropic has shipped source maps in npm.
  • April 7, 2026: Mythos announcement. Wall Street emergency meeting. Glasswing launch.
  • April 8, 2026: A federal appeals court denied Anthropic's stay request in its ongoing Pentagon supply-chain-risk litigation, leaving it excluded from Department of Defense contracts while the case plays out.
  • April 16, 2026 (yesterday): Anthropic released Claude Opus 4.7, which the company described as “less broadly capable than our most powerful model, Claude Mythos Preview.” Axios led with: “Anthropic releases Claude Opus 4.7, concedes it trails unreleased Mythos.” Gizmodo: the release “reads as a promotion for Claude Mythos Preview.”

Marc Andreessen raised the uncomfortable question in Fortune last week: is Mythos being withheld because of security, or because Anthropic lacks the compute to serve it at general-release scale? An industry insider quoted in the New York Post put it more bluntly: “They are trying to deflect from the fact that they can't serve the model because they have no compute.”

On the “$100M Anthropic is committing to Glasswing partners” number that's also making the rounds: per the Glasswing page itself, it's $100M in API usage credits (free API time for partners to use the product they are being asked to validate), plus $4M in direct donations to open-source security organizations. It's a vendor subsidizing customers to generate validation the vendor can then cite. That's a legitimate structure for a customer rollout. It's not a $100M defensive investment in open-source security.

Anthropic's own co-founder on the "special model" framing

Jack Clark — Anthropic co-founder and head of policy — told the Semafor World Economy conference this week:

“It's not a special model. There will be other systems just like this in a few months from other companies. And in a year to a year-and-a-half later, there'll be open-weight models from China that have these capabilities.”

That is the person running policy for the company, publicly stating that the capability driving Mythos's “too dangerous to release” framing is months away from being commoditized by competitors and 12–18 months away from being available in open weights from Chinese labs. Hold that quote next to the Treasury meeting and the CFR piece and the framing stops cohering.

What's actually real

Being honest: Mythos is a capable model. It genuinely chains vulnerabilities better than prior Claudes. It reliably completes multi-step cyber kill-chains that Opus 4.6 could not. It's the first model to solve AISI's 32-step attack range end-to-end. Its pricing — $25/$125 per million tokens on Bedrock, 5× Opus 4.6's — is itself a capability claim Anthropic has to defend against paying customers, and so far the customers aren't publicly refuting it.

The specific techniques are real. The demonstrations are real. The 198 audited reports that agreed with Claude's severity scoring are real.

What isn't real — or more precisely, what the document does not actually claim despite the press coverage — is:

  1. That Mythos is too dangerous for Anthropic's own safety framework to permit release. The system card says, in a footnote, that it isn't.
  2. That Mythos found “thousands of zero-days.” It found 198 audited reports. The thousands are an extrapolation.
  3. That Mythos does something small open-weight models can't also do. Independent researchers have already demonstrated otherwise on several of the specific findings.
  4. That the 27-year-old OpenBSD bug is a remote code execution threat. It appears to be a denial-of-service condition.
  5. That Project Glasswing is a $100M open-source security investment. It's $4M in donations plus $100M in API credits.

The capabilities are real. The framing is strategic. These are not contradictory. But when the framing shapes an emergency Treasury meeting, a White House briefing, and months of regulatory positioning in the run-up to a $60 billion IPO, the distance between them is worth naming.

Six months out from an October IPO target, during an active federal contracting fight, against a compute crunch the competition is publicly mocking, is exactly when an alarming safety narrative is most useful to a company. That doesn't mean the narrative is fake. It means the narrative deserves the same scrutiny as any other IPO-adjacent communication. That scrutiny does not appear to have happened.

Read the footnote.

Sources

Claude Mythos Preview System Card (Anthropic, April 7–8, 2026) · VulnCheck CVE analysis by Patrick Garrity (April 15) · Calif.io MADBugs FreeBSD writeup (March 31) · UK AI Security Institute evaluation (April 14) · Tom's Hardware (April 10) · Fortune interviews with David Lindner and Marc Andreessen (April 13) · Cybernews interview with Marcus Hutchins · Semafor World Economy conference (Jack Clark) · Axios, Bloomberg, Wall Street Journal, CNBC reporting on Anthropic's revenue and compute position (March–April 2026) · CNN, CNBC reporting on Anthropic v. Department of Defense litigation.

Chase Norton is the founder of Redeux Security and an independent security researcher based in Honolulu, HI.

Redeux Security

48-hour adversarial security audits for startups and scale-ups.