OSS Scanner's 88% Claim vs Claude Security
🔍 Comparisons Beginner

OSS Scanner's 88% Claim vs Claude Security

Anthropic’s 88% CVD-bar figure for OSS Scanner is narrower than the free-vs-paid pitch suggests. Here’s what it measures against Claude Security.

The AI Dude · October 10, 2026 · 6 min read

"Eighty-eight percent" is the quality line repeating around Anthropic’s free open-source scanner, and it comes from the company’s Frontier Red Team research post on OSS Scanner: of 97 critical and high findings checked by expert penetration testers, 85 met the bar for Anthropic’s coordinated vulnerability disclosure (CVD) process.

That percentage is doing a lot of work in the free-versus-paid conversation. Put next to Claude Security, it remains a CVD-bar rate on a 97-finding critical and high sample, not a product-wide accuracy score.

Anthropic produced the 88% figure inside a 97-finding expert review

The number did not come from a public leaderboard or from maintainer self-reports alone. Anthropic says it asked the same expert penetration testers who review its CVD findings to check 97 critical and high-severity vulnerabilities from an early OSS Scanner pipeline across 48 projects.

Of those 97, 85 (88%) cleared the CVD bar. Of the remaining 12, Anthropic says 11 were real but duplicated known issues or other findings from the same scan, and only one was invalid as a false positive.

That review sat on top of a larger Glasswing-era backlog the same post describes: over 29,000 candidate vulnerabilities found with recent models across six months of scanning important software, with only about 6,000 manually reviewed and triaged. Maintainers who asked for everything received nearly 5,000 unverified reports with proposed patches. OSS Scanner is the opt-in fast track that ships fully model-generated packs without that human gate.

Partner feedback Anthropic published after several weeks of validating the pipeline also came with names attached. wolfSSL’s Todd Ouska said 74 reports yielded all but two valid findings and five CVEs. OpenSSL Corporation’s Anton Arapov said raw model output was as good as and sometimes better than human reports, particularly when a real exploit was attached. PostgreSQL’s Noah Misch pointed to fixes usable nearly as-is and fast-track access before GA. HotCRP’s Eddie Kohler cited thorough reports with strong permission-model understanding. curl’s Daniel Stenberg said the scanner found multiple issues worth addressing, including one of the worst curl vulnerabilities reported in recent years. Those quotes support demand for speed. They are not the 88% sample.

The percentage only scores critical and high findings against the CVD bar

Read literally, 88% answers one question: among critical and high items already selected for expert review, how many met Anthropic’s internal CVD standard.

It does not score medium or low findings. It does not score the full candidate firehose behind the 29,000 figure. It does not score enterprise private repos under Claude Security’s product framing. The research post draws that product line in one sentence: Claude Security is the general-access code scanning and patching product aimed at enterprises defending their systems, while OSS Scanner gives security audits at no cost to open-source projects that enroll.

OSS Scanner outputs are explicitly fully model-generated, without human review or triage, and Anthropic names its strongest models (including Claude Mythos) as the generators. Each report is described as carrying a self-contained reproducer, an explanation (with bisection when possible), and a candidate patch when available. The 88% check was applied to critical and high items from that style of pipeline, not to every row a maintainer might eventually see in a bulk dump.

On CyberGym, the same post cites a separate academic benchmark arc: LLMs moved from finding under 20% of vulnerabilities early last year to over 85% this year. That 85% is a benchmark recovery rate, not the CVD-bar rate. Collapsing the two into one “about 85–88% accurate” slogan mixes harnesses.

Severity inflation and threat-model misses stay outside the headline rate

Anthropic’s own caveats around the 97-finding review matter more than the single invalid case.

Maintainers, the post says, have seldom called a high or critical finding invalid after receipt. Some have said severity ratings can be inflated, or that the scanner misunderstood the project’s threat model. Those failure modes can leave a finding “real” in a narrow code sense while still wrong for prioritization. An 88% CVD-bar hit rate can coexist with a queue that wastes maintainer time on over-scoped severity or out-of-model risks.

Duplicates are also folded into the non-88% bucket in a way that softens the headline. Eleven of the twelve misses were real-but-duplicate. For disclosure hygiene that is better than inventing bugs. For a small maintainer inbox, a duplicate critical still burns a triage cycle.

What the free path excludes by design is human verification before delivery. Anthropic says it will keep manually disclosing human-verified reports through CVD, especially for projects that lack triage capacity, and that OSS Scanner is the optional fast track for projects that want reports as soon as they exist. Claude Security, in the same framing, is the paid general-access lane for organizations that want scanning and patching inside their own defense workflow rather than an open-source enrollment queue.

Eligibility for the free scanner tracks a short OSS-Fuzz-like bar: critical impact on infrastructure and user security, decided case by case. Core maintainers enroll by submitting a pull request to Anthropic’s GitHub enrollment repo using the standard project template, with more detail in the extended FAQ the post links. That is a different onboarding surface from buying an enterprise scanning product.

  • OSS Scanner: free, opt-in, open-source eligibility, periodic scans, strongest models, no human review on the delivered pack.
  • Claude Security: general-access scanning and patching oriented to enterprise defense, not the no-cost OSS audit lane.

Your triage capacity changes what 88% is worth next to Claude Security

If your project already asked Anthropic for bulk unverified packs, or you run a security-mature maintainer process that can verify reproducers the day they land, the free scanner’s value is speed and volume. The 88% figure then functions as a lower bound on critical/high CVD-worthiness in a reviewed sample, not as a promise that every future pack will match wolfSSL’s 72-of-74 valid ratio.

If you are an enterprise team comparing budget lines, Claude Security is the product Anthropic positions for defending your systems. The research post does not publish Claude Security’s price card, true-positive SLA, or a side-by-side evaluation against the 97-finding OSS Scanner sample. Treating 88% as a transferable enterprise accuracy claim for Claude Security overshoots the evidence. The shared ingredient is model-class capability. The unpaid difference is who owns triage and which codebase class is in scope.

Budget also includes remediation help Anthropic mentions beside the scanner: Claude for OSS offers free Claude Max 20x subscriptions to help remediate vulnerabilities and improve OSS projects, and the Cyber Verification Program opens advanced cyber capabilities to qualifying security professionals. Those are adjacent programs, not substitutes for reading the 88% sample’s boundaries.

For a small OSS team without triage hours, the free pack can still be the wrong default even when the CVD-bar rate looks strong. Anthropic itself keeps the human CVD path for projects that cannot absorb raw model volume. For a funded product security team, paying for Claude Security buys a product lane aimed at enterprise systems. It does not automatically inherit the OSS Scanner validation arithmetic unless Anthropic publishes that transfer study.

The percentage fails as a free-versus-paid accuracy transfer when readers treat Claude Security as covered by the same 97-finding CVD-bar sample, because Anthropic’s post keeps that arithmetic on the OSS Scanner expert review and places Claude Security in the separate enterprise scanning and patching lane.

oss scannerclaude securityanthropicvulnerability scanningcybergym
Share 𝕏 / Twitter Reddit LinkedIn

Keep reading