Skip to content

anthropic

Anthropic's 225 bug finds, one attack in the wild

Claude News

Of 225 vulnerabilities credited to Anthropic and its Project Glasswing, exactly one has been seen in a real attack: a critical SQL injection in the Ghost publishing platform, CVE-2026-26980. The count comes from VulnCheck researcher Patrick Garrity, who has kept a public tracker of these CVEs since April, as The Register reports.

At a glance

  • Garrity's tracker checks every Anthropic-linked CVE against VulnCheck's known exploited vulnerabilities index, and fewer than 0.5 percent of the 225 entries currently carry confirmed exploitation in the wild.
  • Historically, Garrity says, only just under one to two percent of disclosed vulnerabilities are ever weaponised, so this batch is behaving like any other pile of bugs.
  • The harder half sits downstream: in 1Password's study of 6,080 generated patches, frontier models fully resolved the vulnerability 26 percent of the time, leaving remediation to people.

If you have not been following: according to Anthropic's own announcement, Project Glasswing launched on 7 April 2026, giving vetted partner organisations early access to Claude Mythos Preview for defensive security work instead of a public release. The launch partners listed include AWS, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA and Palo Alto Networks, plus over 40 additional organisations that build or maintain critical software infrastructure. Anthropic also committed up to $100M in usage credits and $4M in donations to open-source security organisations.

Only one of the 225 tracked CVEs has been used in an attack

Garrity started the list shortly after the April announcement, and it now runs to 225 vulnerabilities credited to the Anthropic team, to Project Glasswing, or to both. Every entry gets checked against VulnCheck's known exploited vulnerabilities index. As of Monday, one entry had flipped: the Ghost bug. Garrity frames discovery and exploitation as two separate businesses:

There's a big difference between finding vulnerabilities and whether they're actually useful to and will be used by threat actors.

He puts the historical baseline at just under one to two percent of vulnerabilities ever weaponised in the wild. What Anthropic is discovering and disclosing is fairly limited in impact, he told The Register, and is not producing outcomes any different from a random selection of other CVEs.

Anthropic held Mythos Preview back because it finds bugs too well

Anthropic said the model was too risky to release publicly, because its bug-finding and exploitation skills surpass all but the most skilled humans. Access therefore went only to vetted Glasswing participants, who use it defensively, finding and fixing flaws in their own products and in the open source they depend on.

VulnCheck's tracking page attributes dozens of specific CVEs to Anthropic-affiliated researchers working with Claude, including a Firefox batch dated 24 February 2026 running from CVE-2026-2763 to CVE-2026-2799, with CVSS scores from 8.8 to 9.8.

According to Anthropic's announcement, AWS has been testing Claude Mythos Preview inside its own security operations and on critical codebases, and Microsoft ran the model against CTI-REALM, an open-source security benchmark, reporting substantial improvements compared with previous models.

Frontier models produced a clean patch 26 percent of the time

1Password's research team generated and analysed 6,080 patches from two frontier models, OpenAI's ChatGPT-5.5 and Anthropic's Opus 4.8. The fixes fully resolved the vulnerability 26 percent of the time. About 54 percent either failed to resolve it, introduced a new vulnerability, or did both.

According to AI Security Wire, the paper from 1Password's Off-by-1 Labs, titled "Frontier Models' Vulnerability Patches are Often F.L.A.W.E.D.", tested the two models against six recently disclosed open-source flaws, among them an EXIM RCE (CVE-2026-45185) and an ActiveMQ RCE (CVE-2026-34197). Its breakdown adds detail: 20.1 percent of patches fixed the bug but changed application behaviour, and 53.9 percent failed in one way or another.

Veracode measured a different slice. Across more than 100 models and 80 coding tasks, the app security shop found the average security pass rate for AI-generated code was 56 percent.

Why a patch can close the CVE and leave the hole open

The failure mode, according to AI Security Wire, is context: models consistently struggle with patches that require understanding code beyond the lines being changed. The Spring AI SpEL injection case (CVE-2026-22738) is the illustration. SpEL injection can manifest through several expression evaluation contexts, so a patch that seals one leaves the application technically patched against the specific CVE and still exploitable by a variant technique.

Think of a building with four doors keyed alike. Someone reports that the side door opens for anyone, a locksmith changes that one lock, and the ticket closes. The other three doors still open.

That is the gap Garrity points at. The bar for vulnerability discovery is much lower with AI, he says, but coordination, triage, remediation and patch deployment stay largely people-intensive, as Anthropic itself has acknowledged.

These numbers measure what is confirmed rather than what exists: a CVE joins the exploited index once an attack is observed and reported, so one in 225 is a floor, not a census. In our view the interesting mismatch is that Mythos Preview was gated on how well it finds bugs, while every measurement so far puts the weak point in the fixing, which is the expensive, human end of the pipeline.

What would move the 0.5 percent Garrity's tracker keeps running, so the denominator grows with every disclosure credited to Anthropic or Glasswing, while the numerator moves only when somebody catches an attack in the wild. Whether Mythos Preview ever goes generally available has not been stated, and no date for it has been given. The checkpoint worth watching is a second Glasswing CVE joining Ghost in the exploited index, and how long that takes.

Related stories

  1. Claude Fable knocked 20 bits off most popular hashes
  2. A discount Claude reseller was neither cheap nor Claude
  3. Every operation in Anthropic's threat report was disrupted
  4. Reward hacking gaps may have fed Claude's July incidents
  5. Trained to reward hack, Hacker-Opus attacked real targets
  6. ex-OpenAI Researcher Quits Anthropic over AI Safety Fears

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.