Anthropic's bug hunt outruns the people checking it

Over six months, Anthropic's models flagged more than 29,000 candidate vulnerabilities in some of the world's most important open-source projects, but according to Anthropic's own announcement, people have reviewed and triaged only about 6,000 of them. A disclosure dashboard cited by The New Stack shows just 516 vulnerabilities patched upstream as of October 2.
At a glance
- Anthropic has launched OSS Scanner, a free opt-in service that runs periodic scans with its strongest models, including Claude Mythos, and sends unreviewed, fully model-generated reports straight to project maintainers.
- When expert penetration testers checked 97 critical and high-severity findings from an early version across 48 projects, 85 (88%) met the disclosure bar, 11 were real duplicates and one was a false positive.
- Anthropic admits maintainers have flagged inflated severity ratings and misread threat models, and the 516 upstream patches suggest that fixing, not finding, is now the slow part of the pipeline.
If you have not been following, Anthropic used Claude to hunt bugs in open source during what it calls Project Glasswing. It says that on CyberGym, an academic vulnerability-finding benchmark, language models went from finding under 20% of vulnerabilities early last year to over 85% this year. According to Anthropic's disclosure principles, it follows the industry-standard 90-day deadline and paces submissions to what maintainers can actually absorb.
The October 2 dashboard counts 6,157 findings sent and 516 patched upstream
The News Stack pulled the numbers from Anthropic's disclosure dashboard. As of October 2, outside security firms had confirmed 5,674 of the 6,123 findings they reviewed as valid, Anthropic had sent 6,157 findings to maintainers, and 584 CVE and GitHub Security Advisory identifiers had been issued, though some findings received both. Only 516 vulnerabilities had been patched upstream, and roughly 23,000 candidates remain unreviewed.
That snapshot, revision 35, counts 29,439 candidate findings. Disclosures grew from 2,300 across 392 projects on August 26 to 6,157 across 591 projects on October 2, while patches went from 421 to 516. According to Data Studios, 5,103 of the October disclosures were acknowledged by maintainers, and the identifiers split into 219 CVEs and 365 GHSAs.
The figures have drawn scrutiny before. According to Vulnerabilities.ai, VulnCheck's September 10 update found that Anthropic's 91.4% true-positive rate looked considerably lower against at least one documented maintainer account, and that the company's three ledgers did not add up.
Nearly 5,000 unvalidated reports went out because maintainers asked for them
The bottleneck is people. According to The News Stack, the review pipeline relies on six external security research firms to reproduce and triage findings; Data Studios names them as Ada Logics, Anvil, Calif.io, Doyensec, Ophion Security and Trail of Bits.
Anthropic says maintainers who received its first reports increasingly asked for everything, unverified reports with proposed patches included, and it has sent nearly 5,000 of those so far. Its stated reasoning is speed: exploits can now be developed in minutes, so projects that find and fix bugs faster are better placed against attackers racing for the same weaknesses.
Human-verified reports will still go through coordinated vulnerability disclosure (CVD), especially for projects that lack the resources to triage. OSS Scanner turns the bulk route into an optional fast-track, and The News Stack reports it launched as part of Anthropic's broader Cyber Mission. Anthropic cites Google's OSS-Fuzz, which scans open source with fuzzers, as its inspiration, and offers the service free, unlike Claude Security, its enterprise scanning and patching product.
Of 97 critical and high findings checked by pen testers, one was a false positive
To validate an early version, Anthropic asked the expert penetration testers who review its CVD findings to check 97 critical and high-severity vulnerabilities from the scanner across 48 projects. Eighty-five, or 88%, met the CVD bar. Of the remaining 12, 11 were real but duplicated known issues or other findings from the scan, and one was invalid.
Anthropic adds that several weeks of testing with dozens of projects produced hundreds of bug reports, including multiple vulnerabilities it chained into unauthenticated remote code execution exploits. Maintainers, it says, have seldom called a high or critical finding invalid.
Maintainers quoted by Anthropic were positive. Todd Ouska of wolfSSL said all but two of 74 reports were valid and five became CVEs. Daniel Stenberg said the scanner found multiple curl issues, including one of the worst curl vulnerabilities of the last few years. Noah Misch of PostgreSQL said several reports came with fixes usable nearly as-is, and Eddie Kohler of HotCRP praised their grasp of its complex permission model.
Each report ships with a reproducer, a candidate patch and, where possible, a bisection
Each report contains a self-contained reproducer, an explanation of the bug, a candidate patch when available and, where possible, a bisection showing when the bug appeared. Bisection is the string-of-lights trick: test the middle of the project's history, discard the half that behaves, and repeat until a single commit remains.
According to Explainx, a project enrolls through a pull request adding projects/NAME/project.yaml to the anthropics/oss-scanner repository and must supply a Dockerfile that builds it. Setup has network access, but the audit runs offline. Reports go by email to the primary contact, optionally PGP-encrypted, carry no 90-day public disclosure clock and will not be made public.
The packaging is what maintainers single out. Anton Arapov of OpenSSL Corporation, who called early AI reports from about 18 months ago appalling, said Anthropic's reports, raw model output included, matched and sometimes beat human ones, adding:
Particularly when a report comes with a real exploit attached, that's basically job done for an engineer as you can verify it right away
The 88% figure covers 97 critical and high-severity findings from an early scanner version. As The News Stack notes, it says little about lower-severity bugs or performance at scale. Even the 6,157 total mixes 1,333 findings reported after external review with 4,824 sent directly by Anthropic, which may contain false positives. In our view, faster reports fix the wrong end of the pipe: 516 patches suggest maintainers shipping fixes are the slow part.
Whether 516 patches start climbing
Core maintainers of eligible projects can enroll now. Anthropic decides case by case on OSS-Fuzz-like criteria, starting with a "critical impact on infrastructure and user security." It says it will keep refining the scanner from maintainer feedback and as models improve, but it has not said how many projects it will accept or how often scans will run. The next dashboard snapshot will show whether the patch count starts catching up with disclosures.
Related stories
- Anthropic's OSS Scanner skips human review on bug reports
- Claude's workarounds take every Anthropic eval offline
- Cheating on code tests made Anthropic's model sabotage
- One of 225 Anthropic-linked CVEs actually got used
- Claude cheated on 2.4% of its own safety runs
- Self-spreading ideas jump between agents in Anthropic tests
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
