Independent AI intelligence Two editions daily · ET
Fervor AI

Analysis · October 9, 2026 · concept

OSS ScannerAnthropic Cyber MissionClaude Mythosagent-securityfrontier-modelsai-skills

Anthropic's OSS Scanner Will Email You Bugs No Human Has Read

The pilot numbers say the real cost is duplicates, not hallucinations. The threat model file is the only dial you get, so set it before the first report lands.

A frontier lab now sends vulnerability reports to open-source maintainers that no person at the lab has looked at. Anthropic says so in plain words on the Cyber Mission announcement from October 8: "The reports are model-generated and sent without human review."

That sentence would have been a scandal two years ago, when AI bug reports were the thing maintainers begged people to stop sending. An OpenSSL Corporation maintainer quoted in Anthropic's launch post calls the AI reports of about 18 months ago "appalling." So what changed enough that a lab would put its name on unreviewed output, and maintainers would ask for it?

The answer is in the numbers, and the numbers point somewhere most coverage skipped.

What OSS Scanner actually is

OSS Scanner is the open-source half of the Anthropic Cyber Mission. Core maintainers of critical projects can opt in to periodic scans from Anthropic's strongest models, Claude Mythos among them, free of charge. Each report carries a self-contained reproducer, an explanation of the bug (with a bisection to the commit that introduced it where possible), and a candidate patch when one exists.

Projects that don't enroll still get Anthropic's normal coordinated disclosure treatment, where "every report we send generally reflects a finding that a human security researcher has reviewed and confirmed." Enrolling trades that human check for speed. The OSS Scanner page puts it this way: enrolled projects "will receive reports as soon as they're scanned directly from our strongest models."

That trade is the whole product. And the scale explains why Anthropic made it. Over six months, the launch post says, its models produced more than 29,000 candidate vulnerabilities. About 6,000 got manual review. Human triage at that ratio is a queue that never empties.

The number everyone quoted, and the one that matters

Anthropic says it expects "a true-positive rate above 90%." That figure made the headlines. It's also a forward-looking expectation, and it hides the more useful breakdown sitting in the launch post.

In the pilot, expert penetration testers checked 97 critical and high-severity findings across 48 projects. Here's how they landed:

  • 85 met the bar for coordinated disclosure.
  • 11 were real bugs, but duplicates or already known.
  • 1 was a false positive.

Read that again. One fabricated bug in 97. Eleven real ones the maintainer had already seen or would see twice.

The wolfSSL team's experience points the same way: of 74 reports, all but two were valid, and five became CVEs.

So the fear most maintainers carry about AI reports, that they'll spend their weekend chasing a bug that does not exist, is mostly the wrong fear for this tool. The cost OSS Scanner pushes onto a project is triage of a different kind: recognizing a duplicate fast, arguing a severity rating down, and explaining to a model that the "vulnerability" sits outside what the project promises to defend. Anthropic's post admits exactly this, noting that some maintainers said severity ratings can run inflated or that the scanner misread a project's threat model.

That last complaint is the hinge of this whole piece.

You get exactly one dial

Look at what a maintainer can configure. Enrollment is a pull request to anthropics/oss-scanner adding a projects/<project>/project.yaml. The template has three required fields (repo, primary_contact, and a dockerfile the scanner builds your project with) and a handful of optional ones: auto_ccs, homepage, pgp, disabled, and threat_model.

There is no severity threshold. There is no volume cap. The OSS Scanner page mentions neither. You can pause everything with disabled: true or leave entirely by deleting your project directory. Between those two extremes, the only thing shaping what lands in your inbox is the threat model file.

The template calls it "optional but strongly recommended." I'd drop the "optional." The page says the file (by default at .oss-scanner/threat_model.md) can cover scope, adversarial inputs, out-of-scope areas, a severity rubric from critical to low, the report format you want, patch expectations, and deduplication granularity. You can edit it between scans.

That list maps almost one to one onto the pilot's failure modes. Duplicates? Deduplication granularity. Inflated severity? Your rubric. A model that misread what the project defends? Scope and out-of-scope. Anthropic has handed maintainers a steering wheel and labeled it optional.

The clock that isn't there

One more term changes how you should feel about all this. The OSS Scanner page says: "We will not place any form of 90-day coordinated disclosure period on these unvalidated findings."

Compare that with the standard track. Under Anthropic's normal policy, details go public "after 90 days, or after a patch is released, whichever comes first," with a possible 14-day extension, a 7-day target for actively exploited critical bugs, and escalation to an outside coordinator if a maintainer stays silent for 30 days.

Unvalidated fast-track findings carry none of that pressure. If a report later gets validated through the normal process, the 90 days start from when you're told a human validated it. Anthropic does reserve the right to add a disclosure period for some high-severity reports later, with advance notice and an opt-out.

That's a generous arrangement. It also means nobody else is going to force you to triage. A pile of unread emails holding real, reproducible bugs is a worse position than a pile with a deadline, because the deadline at least makes someone open them.

Put this into practice

If you maintain a project that might qualify (the bar mirrors Google's OSS-Fuzz: established projects with real exposure to untrusted input or a large dependent base), here's the lowest-friction order of operations.

1. Write the threat model before you open the enrollment PR. Start from templates/threat_model.md in the oss-scanner repo. Spend most of your time on two sections: out-of-scope (the inputs you never promised to handle safely, the debug-only code paths, the trusted-caller APIs) and the severity rubric. Be concrete. "Memory corruption reachable from a network peer is critical; a crash on a malformed local config file is low" will save you more arguments than any adjective.

2. Make the Dockerfile the environment you actually test in. The scanner builds from it, and every reproducer will run against it. If your real build differs, you'll get findings that only exist in Anthropic's container, and that's a triage cost you created.

3. Decide who reads the email. Reports go to primary_contact. Set a pgp key if you want them encrypted, and know that the template forbids combining pgp with auto_ccs, so an encrypted setup means one inbox. Pick a person, not a list nobody owns.

4. Triage by reproducer first. Run the attached reproducer before you read the explanation. If it doesn't reproduce against your Dockerfile, reply and say so. If it does, check your tracker for duplicates before anything else, since that was the pilot's biggest bucket of non-actionable reports.

5. Feed corrections back into the threat model, not just the email thread. Replying fixes one report. Editing the file shapes the next scan.

Where this falls short

I like this program, and I'd still be careful with it.

The 90 percent figure is an expectation, not a measurement of the production service. The 97-finding pilot was checked by expert testers on critical and high findings only. Nothing published tells you the duplicate rate for medium and low reports, which is where I'd expect the noise to collect.

The eligibility bar leans toward projects that already have capacity. The page says the service is aimed at projects "already able to keep up with verified high/critical vulnerability reports." That makes sense for Anthropic. It also means the small, under-maintained libraries most likely to hide old bugs are the ones least likely to get the fast track.

Delivery is email, for now. The page says the format may change, but today there's no GitHub private advisory flow, so your findings live outside your tracker until someone copies them in.

And the threat model is a document a model interprets. Anthropic's own pilot feedback says the scanner sometimes misread one. A well-written file will help. It won't make the scanner obey it the way a config flag would, so check the first few reports against it and keep editing.

The choice in front of you

Most maintainers will meet OSS Scanner as a link a contributor posts in an issue, asking "should we sign up?" The decision that matters happens right then, before any report arrives.

If you enroll, the reports will mostly be right. Whether they're useful depends almost entirely on a Markdown file you write once and revise after every scan. Write it like the scanner will take it literally, then check whether it did.

Sources: Anthropic Cyber Mission, OSS Scanner launch post, OSS Scanner enrollment page, oss-scanner project template, Anthropic coordinated vulnerability disclosure policy.


Medium metadata

  • Title: Anthropic's OSS Scanner Will Email You Bugs No Human Has Read
  • Subtitle: The pilot numbers say the real cost is duplicates, not hallucinations. The threat model file is the only dial you get, so set it before the first report lands.
  • Tags: Open Source, Cybersecurity, AI, Anthropic, Software Development
  • Canonical URL: set to the fervorai.dev post on import