A vendor tells you its autonomous agent found twenty-one zero-days in one of the most-audited C codebases on the planet, for about a thousand dollars of compute. Your next meeting is in an hour. What can you actually check?
More than you think. The vulnerability record is queryable by anyone, free, without an account. This post gives you the five queries, then runs them on one of 2026's most-cited autonomous-discovery results, on a batch of records that failed the opposite way, and on the exploitation data that says how much any of it should move your patching.
A note on categories before we start
One word covers five different things, so this blog classifies every named incident. A, a jailbreak, no system compromised. B, platform misuse: an external attacker using an AI product against third parties. C, a product vulnerability in shipped software. D, a corporate breach of the vendor's own systems or data. E, model-initiated action during evaluation: a vendor's own model acting against a real third party in a sanctioned test, no attacker involved. Everything named below is Category C or a process change with no incident attached: no A, no B, no D, no E, nobody breached.
The five-minute check: CVE Services, the CVE List, the vendor advisory, the credits field, the exploitation data
Five queries, in order. The first two are the only ones authoritative on whether an identifier exists.
# | Check | Where | Pass looks like | Fail means |
|---|---|---|---|---|
1 | Published CVE record? | Codejavascript | HTTP 200, | HTTP 404, |
2 | Enriched record? | Codejavascript |
|
|
3 | Project acknowledges it? | The project's security page or advisory | Listed with a fixing commit or version | No vendor record of the finding |
4 | Who is credited? | Advisory credits or CNA acknowledgements | Named individuals, or a named entity | No credit, or one that cannot separate tool from human |
5 | Is it exploited? | CISA KEV and exploitation research | The identifier appears in a KEV catalog | Absence is normal, not proof of low risk |
Two cautions. Check 1 tests publication, not reservation: a 404 says no published record exists today, not that an identifier was never reserved. And always run a control: if an identifier does not resolve, query a neighbour from the same numeric block. If the control resolves, the absence is real.
Case study: nine claimed CVE IDs that return 404, and how to run a valid control
The research post 21 Zero-Days in FFmpeg, published by depthfirst, states that its "production autonomous security agent discovered 21 zero-day vulnerabilities in FFmpeg" at "a total cost of roughly $1k (10% of what Anthropic spent using Mythos)". Both numbers are finder-claimed. The post says nine issues were assigned CVEs, CVE-2026-39210 through CVE-2026-39218; the other twelve carry internal DFVULN identifiers.
We ran checks 1 and 2 against all nine on 8 August 2026.
Identifier | CVE Services | NVD REST API |
|---|---|---|
CVE-2026-39210 through CVE-2026-39218 (all nine) | HTTP 404, |
|
CVE-2026-39006 (matched control) | HTTP 200, state |
|
The control is the point. CVE-2026-39006 sits in the same 39xxx block, was assigned by the same CNA, and resolves cleanly from both services on the same day using the same two queries. The block is populated and reachable; the nine claimed identifiers are simply not in it.
Say it precisely, because the loose version is wrong: no published CVE record exists for those nine identifiers as of 8 August 2026. Not "the CVEs are fake", not "the bugs are not real". FFmpeg's own security page publishes CVEs with fixing commit hashes, including entries carrying a finder's internal tracking identifier, so records may land later. What you cannot do is treat "nine assigned CVEs" as verified today. If they publish, those flaws are Category C product vulnerabilities in an open-source media library. Nobody was breached.
Findings are not vulnerabilities and vulnerabilities are not CVEs
Most discovery claims are true narrowly and misleading broadly, because the headline number sits at whichever funnel stage is largest.
Stage | What it means | Illustration from the 2026 record |
|---|---|---|
Raw findings | Anything the system flagged | VulnCheck notes a ledger "claiming Claude had identified 23,019 findings" |
Validated vulnerabilities | Confirmed by a human or by execution | depthfirst claims 21, finder-claimed |
Published CVE records | An identifier that resolves | VulnCheck found 126 of those 23,019 became published CVEs; 0 of depthfirst's 9 resolve today |
Confirmed exploited | Observed in the wild | VulnCheck found 1 of the 126, CVE-2026-26980 in Ghost |
Ask which denominator produced the number. "Six thousand findings" and "six thousand vulnerabilities" differ by an order of magnitude in meaning and not at all in a headline. The buyer question is not "how many did you find" but "how many published, under whose name, and how many did the vendor accept".
The exploitation reality check
The strongest public dataset here is VulnCheck's state of exploitation report for the first half of 2026, published 28 July 2026, which consolidated two AI-attribution datasets and correlated them against VulnCheck KEV. Its finding: "Of 1,061 vulnerabilities attributed to AI-assisted discovery, only 14, or 1.3%, have been confirmed as exploited in the wild, roughly matching the overall exploitation rate of all vulnerabilities in the first six months of the year." For context: 495 KEVs in the period, 23.43 percent exploited on or before CVE publication, median 80 days from publication to KEV.
Two consequences. First, do not apply a severity multiplier for "found by AI" in triage; the evidence does not support one. Second, the same report observed exploitation across 10 of the 28 KEVs it identified in AI systems, including Langflow initial access using CVE-2026-0769 and CVE-2026-5027, then credential harvesting and lateral movement. Those are Category C flaws in AI orchestration software under real exploitation. If your AI stack sits outside pentest scope, close that gap this quarter.
When the record is wrong the other way: nine REJECTED records and the honor system
Verification cuts both ways. Nine records assigned by MITRE, CVE-2026-51296 through CVE-2026-51304, were published on 27 July 2026 (two carry no publication date) and moved to REJECTED on 31 July 2026: claims of Category C flaws, withdrawn. Each now carries the same reason, retrievable from CVE Services today: "DO NOT USE THIS CVE RECORD. ConsultIDs: none. Reason: This record was withdrawn by its CNA. Further investigation showed that it was not a security issue. Notes: none."
The mechanism was described the same day on the oss-security list by Alan Coopersmith of Oracle, in a message timestamped Fri, 31 Jul 2026 18:03:40 -0700 and archived under 1 August 2026. Quote both dates, or a checker will think one is wrong. His conclusion is the line to read to your team: MITRE and most other CNAs that assign CVEs for code they do not produce themselves "operate on the honor system, and trust CVE requesters to have verified the information they provide, since the CNA is often not in a position of being able to verify the report themselves".
A CVE identifier is therefore not evidence that a vulnerability exists, only that somebody asserted one and a CNA accepted it. That is why check 1 reads the record's state, not merely whether a page loads.
curl in January versus curl in April
The most-quoted evidence that AI reporting is a net negative is curl ending its bug bounty. That is real, and curl's own numbers overtook it inside three months.
26 January 2026 | 22 April 2026 | |
|---|---|---|
Source | ||
Confirmed rate | "north of 15%" historically, "below 5%" from 2025 | "back to and even surpassing the 2024 pre-AI level, meaning somewhere in the 15-16% range" |
Volume | Rising, with an "explosion in AI slop reports" | "about double the rate we had through 2025, which already was more than double from previous years" |
Verdict on slop | "Not even one in twenty was real" | "The slop situation is not a problem anymore" |
The variable that changed was not AI. In between, curl moved reporting to GitHub, called that a mistake and returned to its previous platform on 1 March 2026, reward still removed. Stenberg's April read is that "almost every security report now uses AI to various degrees" but that "they are mostly very high quality". If someone cites the January post at you now, they are citing a snapshot its author superseded.
What maintainers actually changed: named-human verification, not prohibition
"Open source is banning AI bug reports" is not what the primary sources say. They show a shift to attributable human accountability.
FFmpeg states that "Automated submissions are not accepted" and warns of "a spike in AI generated, false positives", asking reporters to "make sure that what you report are real issues by careful human verification". The requirements are the operative part: the "name(s) or alias(es) of the human reviewer(s) who verified the report", a reproducible testcase, a git commit hash, and stack traces. Machine assistance is permitted; anonymous machine output is not.
GNOME ran both sides in public. On 8 June 2026 Michael Catanzaro published Please Do Not Ban AI-Assisted Issue Reports, arguing quality has become bimodal: reports are "now often better or worse than before", better when an experienced human works with a good tool and worse when an inexperienced one does not. On 20 July 2026 the same maintainer cut GNOME's disclosure deadline from 90 days to 30 for issues reported on or after 1 August 2026, opening with "due to the increase in AI-generated security vulnerability reports" while also noting "the shorter deadline would probably work better for GNOME even if not for the increase in AI-generated issue reports". Both statements sit in the same post.
OpenSSF's vulnerability disclosures working group has had an open issue since 4 February 2026, AI-SLOP: Develop best current practises for Open Source maintainers, whose stated goal is explicit: "The goal is to reduce slop, not ban AI entirely." Its principles are human-in-the-loop accountability, disclosure of substantial AI assistance, no autonomous agents opening pull requests, and an unchanged quality bar.
Note what this does to check 4. The Redis advisory covering CVE-2026-23479 and four siblings, all Category C, has a section headed "Who gets the credit?" recording that CVE-2026-23479 was "reported by independent researchers Team Xint Code (Tim Becker @tjbecker, Jacob Newman, and Juno IM)". It names no AI system anywhere. If you have read this one described as an autonomous AI discovery, the vendor record does not say so.
Questions for your AI security vendor questionnaire
Ask | A good answer | A red flag |
|---|---|---|
Which denominator is your headline number? | Names the stage: raw findings, validated, published, exploited | Uses "findings" and "vulnerabilities" interchangeably |
List identifiers we can verify ourselves | CVE or GHSA IDs that resolve, plus the vendor advisory | Internal tracking IDs, or IDs with no published record |
Who is named in the credits field? | Named humans, or an entity matching the advisory | "Our AI", with no record trace |
What is your false-positive rate, and measured by whom? | Independent assessment, method stated | Self-assessed, no method |
Does a human validate a finding before it reaches us? | Yes, named role, at a defined severity threshold | "The model verifies it" |
Can you reproduce a finding on demand mid-engagement? | Yes, with a reproducible testcase | Narrative write-up only |
The same discipline belongs in scope, not only in procurement. A 2026 penetration test should treat the AI stack as ordinary application and infrastructure surface, because that is where the observed exploitation sits: orchestration and workflow platforms, inference endpoints, the RAG data plane, and the CI runners that execute agent output.
That is how Stingrai structures its own work. Snipe, our autonomous agent, covers web application testing, and its findings carry reproducible evidence rather than narrative assertions. On our Hybrid tier a human penetration tester validates every finding before it reaches you; our testers hold CREST CRT and Stingrai is a CREST-accredited penetration testing service provider at firm level. Our team has 18 published CVEs, so we have been on the reporting side of this process. Web application tiers are Autonomous at US$3,000 one-time or US$450 per month and Hybrid at US$6,800 one-time or US$1,275 per month, both carrying a "No High or Critical Finding = Don't Pay" guarantee.
Frequently Asked Questions
How do I check if a CVE is real?
Query the authoritative record at https://cveawg.mitre.org/api/cve/ followed by the identifier and read the state field: a published record returns HTTP 200 with a state of PUBLISHED, and an identifier with no published record returns HTTP 404 with the error CVE_RECORD_DNE. Cross-check the enriched record at https://services.nvd.nist.gov/rest/json/cves/2.0?cveId= followed by the identifier, where totalResults of 1 means it exists and 0 means it does not. Then confirm the project's advisory lists it with a fixing commit, and run a control query on a neighbouring identifier so a missing record is not mistaken for an outage.
Did an AI agent really find 21 zero-days in FFmpeg?
The 21-bug count and the roughly US$1,000 compute figure come from the finder's own research post and have not been independently reproduced. That post names nine CVE identifiers, CVE-2026-39210 through CVE-2026-39218, and as of 8 August 2026 none returns a published record from CVE Services or the NVD REST API, while a matched control in the same block, CVE-2026-39006, resolves from both the same day. The precise statement is that no published CVE record exists for those nine identifiers as of that date; reservation status is not publicly queryable, so it does not mean they were never reserved.
Are AI-generated vulnerability reports accurate?
Quality is bimodal rather than uniformly poor. curl's Daniel Stenberg reported the confirmed-vulnerability rate falling below 5 percent in 2025, then on 22 April 2026 reported it back in the 15 to 16 percent range at roughly double the volume, writing that the slop situation was no longer a problem. GNOME maintainer Michael Catanzaro describes reports as now often better or worse than before, depending on the experience of the human behind them. The discriminator is not whether AI was involved but whether a named human verified and can reproduce the finding.
Why did MITRE reject those CVEs?
Nine records, CVE-2026-51296 through CVE-2026-51304, were published on 27 July 2026 and moved to REJECTED on 31 July 2026. Each carries the same reason, retrievable from CVE Services: the record was withdrawn by its CNA after further investigation showed that it was not a security issue. On the oss-security list the same day, Alan Coopersmith of Oracle noted that MITRE and most CNAs assigning CVEs for code they do not produce operate on the honor system, trusting requesters to have verified what they submit, which is why a CVE identifier alone is not proof that a vulnerability exists.
Did curl really end its bug bounty because of AI?
Yes, and the story did not stop there. curl closed its bug bounty on 31 January 2026 after the confirmed-vulnerability rate fell from north of 15 percent to below 5 percent, with Stenberg citing an explosion in AI slop reports. It then moved reporting to GitHub, called that a mistake, and returned to its previous platform on 1 March 2026 with the reward still removed. By 22 April 2026 Stenberg reported roughly double the 2025 volume at a confirmed rate back in the 15 to 16 percent range, so citing only the January post now quotes a snapshot its author superseded.
What percentage of AI-found vulnerabilities are actually exploited?
VulnCheck's state of exploitation report for the first half of 2026, published 28 July 2026, consolidated two AI-attribution datasets and correlated them with VulnCheck KEV. Of 1,061 vulnerabilities attributed to AI-assisted discovery, 14, or 1.3 percent, were confirmed exploited in the wild, which the report describes as roughly matching the overall exploitation rate for all vulnerabilities in that period. Its conclusion is that the data does not suggest AI-discovered vulnerabilities are inherently more likely to be exploited than those found traditionally, so no severity multiplier for AI attribution is justified in triage.
Do open source projects ban AI-generated bug reports?
Some ban AI-generated content broadly, but the security-specific policies are narrower than the headline suggests. FFmpeg states that automated submissions are not accepted while requiring the names or aliases of the human reviewers who verified the report, which permits machine assistance behind a named human. GNOME's security maintainer publicly argued against banning AI-assisted issue reports, and OpenSSF's open working-group issue states its goal as reducing slop rather than banning AI entirely. The consistent requirement is attributable human verification, not prohibition.
What should I ask an AI penetration testing vendor about their findings?
Ask which denominator the headline number uses: raw findings, validated vulnerabilities, published identifiers, or confirmed exploited. Ask for CVE or GHSA identifiers you can resolve yourself, for the names in the credits field, and for how many submissions affected vendors accepted versus rejected. Ask whether a human validates every finding before delivery, whether findings can be reproduced on demand during the engagement, and what happens when one cannot. A vendor answering all of these with verifiable specifics is describing a programme; one that cannot is describing a slide.
References
depthfirst, 21 Zero-Days in FFmpeg: https://depthfirst.com/research/21-zero-days-in-ffmpeg
CVE Services record API (queried 8 August 2026 for CVE-2026-39209 through CVE-2026-39219, CVE-2026-39006, and CVE-2026-51296 through CVE-2026-51304): https://cveawg.mitre.org/api/cve/CVE-2026-51302
NVD REST API (queried 8 August 2026): https://services.nvd.nist.gov/rest/json/cves/2.0?cveId=CVE-2026-39006
Alan Coopersmith, Rejected CVE reports against SQLite, libraw, ESP32-audioI2S, oss-security (timestamped 31 July 2026, archived 1 August 2026): https://www.openwall.com/lists/oss-security/2026/08/01/2
VulnCheck, State of Exploitation, first half 2026 (28 July 2026): https://www.vulncheck.com/blog/state-of-exploitation-1h-2026
Daniel Stenberg, The end of the curl bug-bounty (26 January 2026): https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/
Daniel Stenberg, curl security moves again (25 February 2026): https://daniel.haxx.se/blog/2026/02/25/curl-security-moves-again/
Daniel Stenberg, High-Quality Chaos (22 April 2026): https://daniel.haxx.se/blog/2026/04/22/high-quality-chaos/
FFmpeg Security policy and CVE list: https://www.ffmpeg.org/security.html
OpenSSF, AI-SLOP: Develop best current practises for Open Source maintainers (opened 4 February 2026): https://github.com/ossf/wg-vulnerability-disclosures/issues/178
Michael Catanzaro, Please Do Not Ban AI-Assisted Issue Reports (8 June 2026): https://blogs.gnome.org/mcatanzaro/2026/06/08/please-do-not-ban-ai-assisted-issue-reports/
Michael Catanzaro, Some Changes to GNOME Security Tracking (20 July 2026): https://blogs.gnome.org/mcatanzaro/2026/07/20/some-changes-to-gnome-security-tracking/
Redis security advisory for CVE-2026-23479 and siblings: https://redis.io/blog/security-advisory-cve202623479-cve202625243-cve-2026-25588-cve202625589-cve-2026-23631/
NVD records for CVE-2026-0769, CVE-2026-5027 and CVE-2026-26980 (queried 8 August 2026): https://services.nvd.nist.gov/rest/json/cves/2.0?cveId=CVE-2026-0769



