main logo icon

Published on

July 23, 2026

|

13 min read

AI-Attributed CVEs of 2026: A Primary-Sourced Tracker With Verified Vendor Credits

A running, primary-sourced tracker of the CVEs that autonomous AI agents and AI security firms actually discovered in 2026. Each row is keyed to its finder, credit bucket, CVSS, fixed versions, and primary advisory.

Arafat Afzalzada

Arafat Afzalzada

Founder

AdvisoriesLLM Security

Summarize with AI

ChatGPTPerplexityGeminiGrokClaude

TL;DR

In 2026, AI stopped being a lab demo in vulnerability research and started earning real CVE credits. This tracker records the confirmed cases, keyed on CVE id and, more importantly, on who actually found the bug. An autonomous agent (XBOW) was credited on a critical CVSS 9.8 remote code execution flaw in a Microsoft cloud service (CVE-2026-21536). Google's Big Sleep agent found a SQLite flaw (CVE-2025-6965) and helped cut off an in-the-wild exploit before it landed. An AI lab, depthfirst, reported 21 FFmpeg zero-days for roughly US$1,000 in compute, with nine already numbered (CVE-2026-39210 through CVE-2026-39218). Apple credited the research firm Calif.io, working with Claude and Anthropic Research, on a macOS flaw (CVE-2026-28952). And a critical Exim RCE (CVE-2026-45185) is the corrective case: it was discovered and reported by a human, XBOW Security Lab head Federico Kirschbaum, and the AI only entered afterwards, in an internal race to write the exploit that the human won, with the autonomous system landing a working exploit only against a deliberately weakened target, so it carries no AI finder credit at all. The differentiator here is accuracy: getting the finder and the credit bucket right, with every claim traced back to its own advisory. One often-repeated pairing did not make the tracker because the attribution did not hold up. The operational takeaway for defenders is continuous validation, because the discovery cadence is now measured in agent runs, not annual engagements.

An autonomous AI agent was credited on a critical CVSS 9.8 remote code execution flaw in a Microsoft cloud service in March 2026, a first for a machine finder on a bug that severe (NVD, CVE-2026-21536; XBOW, 2026). Two months later an AI lab reported 21 zero-days in FFmpeg for roughly US$1,000 of compute, with nine already assigned CVE identifiers (depthfirst, 2026). A year earlier, Google's Big Sleep agent had already found a SQLite flaw and helped cut off an in-the-wild exploit before it landed (Google, 2025). The AI attacker has stopped being a conference slide and started showing up in the CVE record itself.

That is exactly why attribution now matters more than the headline. Every one of these claims travels through press releases and secondary write-ups that blur who found what, and the difference between "an autonomous agent found it" and "a human used a model as a tool" is the whole story for a defender trying to model the threat. This tracker exists to keep that record straight.

The direct answer

Which real CVEs did autonomous AI agents and AI security firms discover in 2026, who actually found them, and are they patched? Five verified CVE records across 2025 and 2026 sit in this tracker, four of them carrying an AI finder credit, and all of them are fixed. Three were led by autonomous agents: a Microsoft cloud RCE (XBOW), a SQLite memory-corruption flaw (Google Big Sleep), and a cluster of FFmpeg parser bugs (depthfirst). One was an AI-assisted collaboration where a research firm worked with Claude, credited by Apple on a macOS flaw. The fifth is here precisely because it is not an AI find: the critical Exim RCE was discovered and reported by a human, XBOW Security Lab head Federico Kirschbaum, and the AI only entered afterwards, in an internal exploit-writing race the human won, with the autonomous system succeeding only against a deliberately weakened target. One widely repeated pairing was left off because the attribution did not survive a check against the primary advisory.

Ai Cve Tracker Matrix V3

TL;DR: the numbers that matter

  • Highest-severity AI-credited CVE (Mar 2026): an autonomous agent was credited on a CVSS 9.8 Microsoft Devices Pricing Program RCE, CVE-2026-21536 (NVD).

  • Cheapest large haul (Jun 2026): depthfirst's agent reported 21 FFmpeg zero-days for about US$1,000, nine now numbered CVE-2026-39210 through -39218 (depthfirst, 2026).

  • First AI-foiled exploit (Jul 2025): Google's Big Sleep found SQLite flaw CVE-2025-6965 and helped stop an in-the-wild exploit before it landed (Google, 2025).

  • A vendor names the model (May 2026): Apple credited Calif.io, in collaboration with Claude and Anthropic Research, on macOS flaw CVE-2026-28952 (Apple, 2026).

  • Critical, but human-found (May 2026): the Exim "Dead.Letter" RCE CVE-2026-45185 (CVSS 9.8) was discovered and reported on 1 May 2026 by Federico Kirschbaum, who heads XBOW's Security Lab, and the human then won the internal race to write a working exploit, with the autonomous system succeeding only against a deliberately weakened target (NVD; BleepingComputer, 2026; The Hacker News, 2026).

  • The oldest bug an agent found: one FFmpeg stack overflow in the cluster had been latent since 2003, roughly 23 years, before an AI agent surfaced it (The Hacker News, 2026).

  • Every entry is patched: all five records point to a fixed version or a server-side mitigation, so this is a patch-status tracker, not an open-threat feed.

  • One exclusion on accuracy: the Kimi K3 and Redis pairing was left off because that specific claim has no CVE assigned yet, while the numbered Redis CVEs carry named human credits and CVE-2026-23479 belongs in the AI-assisted bucket rather than the autonomous one.

Key takeaways

  • The credit bucket is the whole story, and most coverage gets it wrong. "Autonomous," "AI-assisted," and "human-found with an AI racing behind it" are three different threat claims. The Exim RCE is the cleanest example of the third: a human found and reported it, and a human also won the exploit-writing race that followed, with the autonomous system landing a working exploit only against a deliberately weakened target. The line most quoted from that episode belongs to the researcher rather than the company, and Kirschbaum said he does not think "LLMs alone are quite ready to write exploits against real-world software" yet (BleepingComputer, 2026; The Hacker News, 2026).

  • Autonomous agents now reach severity that used to require elite humans. A machine finder credited on a CVSS 9.8 RCE (NVD, CVE-2026-21536) is a different world from the "AI finds low-hanging XSS" framing of two years ago.

  • The economics changed before the tooling matured. Roughly US$1,000 of compute for 21 FFmpeg zero-days (depthfirst, 2026) means an attacker can now afford to point an agent at every dependency you ship, not just your crown-jewel app.

  • Old code is the new frontier. A 23-year-old FFmpeg stack overflow (The Hacker News, 2026) shows agents are strongest where humans stopped looking: mature parsers, demuxers, and libraries buried deep in the software supply chain.

  • Defense has to move at agent cadence. When discovery is measured in runs rather than annual engagements, point-in-time testing leaves a widening gap. Continuous validation is the honest response, and it is why we fold autonomous coverage into our AI-driven offensive security operations.

Methodology

This is a running tracker, refreshed as new dated, verifiable AI-attributed CVEs publish. Inclusion has a hard bar: a real CVE identifier, a named finder, a credit that survives a check against the primary advisory, and a confirmed fixed version or mitigation. Sources used for this pass, with dates:

  • Apple. About the security content of macOS Tahoe 26.5, published May 2026, plus the matching NVD record for CVE-2026-28952.

  • NVD (NIST National Vulnerability Database). Per-CVE records for CVE-2026-21536, CVE-2026-45185, CVE-2025-6965, and the reserved FFmpeg cluster CVE-2026-39210 through CVE-2026-39218.

  • XBOW. The vendor's own write-ups on the Microsoft RCE cluster (March 2026) and the Exim "Dead.Letter" RCE (May 2026). Note that xbow.com returned HTTP 429 to every automated request on our verification passes, so the Exim entry is corroborated here by reachable independent reporting rather than resting on the vendor link alone.

  • BleepingComputer and The Hacker News. Independent May 2026 reporting on CVE-2026-45185, used to confirm the finder credit and disclosure date for that row.

  • Google. Cloud CISO Perspectives on the Big Sleep agent and SQLite CVE-2025-6965 (2025).

  • depthfirst. The lab's research write-up on 21 FFmpeg zero-days (June 2026), with per-CVE confirmation via NVD.

The research cutoff for this pass is July 2026. CVSS scores are the CVSS 3.1 base values from NVD, noted per row; where a publisher used CVSS 4.0, both are given in the entry detail. Where a CVE cluster is still reserved at NVD, its severities are marked pending rather than estimated, because inventing a score would defeat the point of the tracker. Any figure that could not be reached on at least one verification pass against a named primary source was dropped rather than printed, and no proof-of-concept code or exploit steps are reproduced here. Every figure below links back to its primary advisory so any claim can be audited inline. For the broader disclosure and exploitation context around these individual entries, see our 2026 vulnerability statistics.

The 2026 AI-attributed CVE tracker

This is the running table. Each row states the affected and fixed versions, the finder, the credit bucket (the finder type), the CVSS 3.1 base score, and the single most useful "what to patch now" line. The CVE identifier links to the primary advisory.

CVE (advisory)

Affected, fixed

Finder

Credit bucket

CVSS 3.1

What to patch now

CVE-2026-21536

Microsoft Devices Pricing Program (cloud); mitigated server-side

XBOW (New York, US)

Autonomous AI agent

9.8 Critical

No customer action; confirm your March 2026 Patch Tuesday rollout for the rest of the bulletin

CVE-2026-45185

Exim 4.97 to 4.99.2, GnuTLS builds; fixed in 4.99.3

XBOW (Federico Kirschbaum, Security Lab)

Human-found, AI-raced (see row note)

9.8 Critical

Upgrade Exim to 4.99.3; OpenSSL-only builds are not affected

CVE-2025-6965

SQLite before 3.50.2; fixed in 3.50.2

Google Big Sleep (DeepMind, Project Zero)

Autonomous AI agent

7.7 (7.2 under CVSS 4.0)

Update SQLite to 3.50.2 or later, and rebuild the many apps that embed it

CVE-2026-39210 to -39218

FFmpeg parsers and demuxers; fixed upstream

depthfirst (San Francisco)

Autonomous AI agent

Reserved at NVD (pending)

Update FFmpeg and any bundled copies; prioritize services that ingest untrusted RTSP or AV1-over-RTP

CVE-2026-28952

macOS Tahoe before 26.5, iOS and iPadOS before 18.7.9, Sequoia before 15.7.7, Sonoma before 14.8.7

Calif.io with Claude and Anthropic Research

AI-assisted (human plus AI)

7.5 High

Update to the fixed Apple OS builds

Row note, CVE-2026-45185. This row carries a human finder credit, not an AI one. Federico Kirschbaum, who heads XBOW's Security Lab, discovered and reported the flaw on 1 May 2026, and he is XBOW's own staff rather than an outside researcher (BleepingComputer, 2026; The Hacker News, 2026; NVD). The AI angle came only in the exploit-writing phase that followed, an internal contest in which XBOW's autonomous system raced XBOW's own humans, who used a model as an assistant. The human won, but the machine did not come away empty-handed: XBOW Native produced a working exploit against a simplified target, an Exim server with no ASLR and a non-PIE binary, and on a second attempt against a machine with ASLR though the binary was still non-PIE, going after Exim's own allocator rather than off-the-shelf glibc allocator techniques (XBOW, 2026). Two things follow, and the tracker states both plainly rather than softening them: on this entry the AI lost the discovery step outright and cleared the last mile only against a target that had been weakened for it, and the vendor write-up is currently the weakest of the four citations here, because xbow.com returned HTTP 429 to every automated request on our verification passes, which is why the two trade reports above carry the load. The much-quoted "not ready" line from this episode is the researcher's personal assessment, not a corporate verdict, and it is quoted in full in the entry detail below.

Reading the credits: three buckets, not one

The reason this tracker leads with the finder rather than the flaw is that the three credit buckets describe three genuinely different threats. Collapsing them into "AI found a bug" is where most reporting goes wrong, and it is where a defender loses the plot.

Ai Cve Tracker Buckets V4

Autonomous AI agent means an agent drove the discovery end to end, with a human triaging before disclosure rather than doing the finding. XBOW on the Microsoft RCE, Big Sleep on SQLite, and depthfirst on the FFmpeg cluster all sit here. This is the bucket that should change your threat model, because it scales with compute, not headcount.

AI-assisted, human plus AI means a human researcher used a model as a power tool. The Apple macOS credit names both sides of that collaboration, and that phrasing is deliberate. The find is real and the model mattered, but a skilled human was still steering.

Human-found, AI-raced is the one that gets mislabeled most, because the finding firm is an AI security company and that alone is enough for coverage to file it under "AI found it." The Exim RCE belongs here. A person found and reported it: Federico Kirschbaum, head of XBOW's Security Lab, on 1 May 2026 (BleepingComputer, 2026; The Hacker News, 2026). The AI showed up one step later, when XBOW's autonomous system raced its own humans to write a working exploit and the humans won, with the machine succeeding only against a deliberately weakened target (XBOW, 2026). The line most often quoted from that episode belongs to the researcher, not the company: Kirschbaum said he does not think "LLMs alone are quite ready to write exploits against real-world software" yet (BleepingComputer, 2026). Logging this row as autonomous, or even as a shared find, would overstate the state of the art twice over.

The entries in detail

CVE-2026-21536: an autonomous agent on a Microsoft 9.8

The highest-severity AI-attributed credit of the year went to XBOW, a New York based autonomous security company, for a remote code execution flaw in the Microsoft Devices Pricing Program. NVD records it as an unrestricted file upload weakness (CWE-434) with a CVSS 3.1 base score of 9.8, allowing an unauthenticated attacker to execute code with no user interaction (NVD, CVE-2026-21536). Microsoft disclosed it in the March 2026 Patch Tuesday cycle and, because the affected component is delivered as a cloud service, mitigated it server-side (Krebs on Security, 2026).

For defenders, the significance is not a patch action, since there is nothing to install for this specific CVE. It is the precedent. An autonomous system found and reported a critical, unauthenticated RCE in a major vendor's cloud surface without source access. The right response is to confirm your broader March 2026 Patch Tuesday rollout landed, and to assume the same class of agent is now enumerating your own internet-facing upload paths.

CVE-2025-6965: Big Sleep and the exploit that never landed

Google's Big Sleep, built by DeepMind and Project Zero, found a memory-corruption flaw in SQLite affecting all versions before 3.50.2, fixed in 3.50.2 (NVD, CVE-2025-6965). NVD scores it 7.7 under CVSS 3.1, and Google scored it 7.2 under CVSS 4.0. The standout detail is operational: Google says this was the first time an AI agent was used to directly foil efforts to exploit a vulnerability in the wild, combining threat intelligence with the agent to catch a bug that was "known only to threat actors" before it could be weaponized (Google, 2025).

SQLite is one of the most widely embedded pieces of software on earth, so the patch action is bigger than it looks. Updating your system SQLite is the easy part. The harder part is the long tail of applications, mobile apps, and appliances that bundle their own copy and need a rebuild to pick up 3.50.2.

CVE-2026-39210 to -39218: 21 FFmpeg zero-days for the price of a laptop

In June 2026, the AI lab depthfirst reported that its autonomous agent had scanned roughly 1.5 million lines of FFmpeg C code and produced 21 confirmed zero-days, at a run cost of about US$1,000. Nine have CVE identifiers so far, CVE-2026-39210 through CVE-2026-39218, and the rest are fixed upstream while awaiting numbering (depthfirst, 2026). The bugs are mostly heap or stack overflows in parsers and demuxers, spanning the TS demuxer, the VP9 decoder, the RTP depacketizer, and more, with one stack overflow in service-description-table handling that had sat untouched since 2003 (The Hacker News, 2026). These NVD records are reserved as of this pass, so their individual CVSS scores are still pending, and this tracker will not guess them.

The patch guidance is blunt. Pull the fixed upstream FFmpeg build or your distribution's security update as soon as it lands, and remember that embedded copies of FFmpeg inside media pipelines, transcoders, and device firmware need patching too. Prioritize anything that ingests untrusted RTSP or AV1-over-RTP, which is where the most serious of these live. The strategic lesson is that a US$1,000 agent run can now sweep a dependency that decades of manual review and fuzzing had not fully cleared.

CVE-2026-28952: Apple names the model in the credit

Apple's macOS Tahoe 26.5 advisory credits Calif.io, in collaboration with Claude and Anthropic Research, for an integer overflow that a malicious app could use to cause unexpected system termination (Apple, 2026). NVD lists a CVSS 3.1 base score of 7.5 and confirms the fix shipped across macOS Tahoe 26.5, macOS Sequoia 15.7.7, macOS Sonoma 14.8.7, and iOS and iPadOS 18.7.9 (NVD, CVE-2026-28952). This is the AI-assisted bucket in its clearest form. A research firm used the model as part of its workflow, and the vendor's acknowledgement names both the firm and the AI, rather than pretending either did it alone.

The patch action is simple and should already be in your fleet management queue: update to the fixed Apple OS builds. The wider signal is that a mainstream vendor's security acknowledgements now list an AI collaborator by name, which normalizes AI-assisted research as a standard, creditable method rather than a novelty.

CVE-2026-45185: the Exim RCE a human found, and the AI only raced

The Exim "Dead.Letter" flaw is a remotely reachable use-after-free (CWE-416) in the BDAT body-parsing path, affecting Exim builds from 4.97 up to 4.99.2 that use GnuTLS, fixed in 4.99.3. NVD scores it 9.8 under CVSS 3.1, an unauthenticated network attacker path to arbitrary code execution (NVD, CVE-2026-45185). It is in the tracker without an AI finder credit, and the reason is the most instructive thing on this page.

A human found it. Federico Kirschbaum, head of the Security Lab at XBOW, is credited with discovering and reporting the flaw, on 1 May 2026 (BleepingComputer, 2026; The Hacker News, 2026). He is XBOW's own staff, not an outside researcher who happened to beat a vendor's agent, which is how this story is usually retold. The AI entered only in the phase that followed, when turning the finding into a working exploit became an internal contest between XBOW's autonomous system and XBOW's own humans working with a model as an assistant. The humans won that too, though the machine did not come away empty-handed, and that half matters on a page about precision. XBOW Native produced a working exploit, but only against a simplified target: an Exim server with no ASLR and a non-PIE binary. On a second attempt it achieved an exploit against a machine with ASLR, though the binary was still non-PIE. It also went after Exim's own allocator rather than reaching for off-the-shelf glibc allocator techniques (XBOW, 2026). That is a genuine result, and it is also a long way from a hardened production mail server, which is the only target that matters to a defender. Standard binary hardening was doing real work in that gap.

The line most repeated from this episode also deserves its correct owner. It is the researcher's personal assessment, not a corporate verdict: Kirschbaum said he does not think "LLMs alone are quite ready to write exploits against real-world software" yet (BleepingComputer, 2026). He was equally clear about the other side of it, crediting AI tools with a crucial role in helping humans understand unfamiliar code and investigate suspicious areas far faster than they otherwise could. Both halves are the finding: models did not close the last mile unaided against a realistic target, and they measurably sped up the person who did.

State the scoreboard without softening, because it is the strongest evidence the tracker has: on the highest-severity entry an AI security firm put its name to this year, the machine lost the discovery step outright, and it cleared the last mile only after the target had been weakened for it. Note also that xbow.com returned HTTP 429 to every automated request on our verification passes, so the vendor write-up is the least reachable of these citations, and the two trade reports are what make this row auditable by a reader or a crawler.

Patch action: upgrade Exim to 4.99.3. Builds compiled against OpenSSL rather than GnuTLS are not affected, but given Exim's share of internet mail, GnuTLS builds should be treated as urgent.

What did not make the tracker, and why

Accuracy means saying no as often as yes. The most common pairing we were asked about was Kimi K3 and Redis. It is not in the table for two independent reasons. First, the Redis advisory CVEs that circulate alongside that story carry named human credits, and the one with a genuine AI angle, CVE-2026-23479, belongs in the AI-assisted bucket rather than the autonomous one, because Team Xint Code found it by running Theori's autonomous code auditing tool Xint Code. That AI-assist detail comes from the finder's own tool disclosure rather than from the Redis advisory, which never mentions AI, which is why it is noted here rather than given a tracker row. Second, the separate, widely shared claim that Kimi K3 found a Redis issue in minutes has no CVE assigned to it, so it fails the tracker's basic bar of a real identifier with a verifiable finder. A model doing something impressive in a demo is not the same as an AI holding a credit on a numbered, fixed vulnerability, and this tracker only records the latter. If a Redis CVE is later credited to an autonomous agent, it earns a row then, not before.

Ai Cve Tracker Timeline V3

What this means for defenders

The through-line of the 2026 record is not that AI replaced human researchers. It is that the discovery side of the industry now runs at a cadence and cost that point-in-time defense was never designed for. When an agent can sweep a dependency for zero-days at the price of a laptop, the question is not whether your annual pentest was thorough. It is what has changed in the eleven months since it finished. Our take on that shift, and how an autonomous attacker actually operates, is laid out in Inside Agentic Red Teaming and in our read of the AI offensive security tool boom.

For teams turning this into a plan, three moves follow directly from the tracker:

  • Match the attacker's cadence with continuous validation. The classes AI is strongest at, mature libraries and internet-facing services, are exactly the ones that drift between annual tests. Testing continuously is how you close the gap between disclosure and detection.

  • Route the work to the right tester by vulnerability class. The complex web application and API authorization layer, where broken access control, IDOR, and business-logic flaws live, is where Stingrai's Snipe operates as an autonomous agent, trained to hunt exactly those high-impact classes rather than stopping at known-class noise. The datastore, media library, mail server, and OS kernel classes on this tracker, the SQLite, FFmpeg, Exim, and macOS style bugs, route to senior human pentesters through our network and infrastructure testing and red teaming practices. Getting that split right is the same discipline as getting the credit bucket right.

  • Keep humans on the discovery step and the last mile. The Exim entry is the proof, and it cuts harder than the usual version of the story. A human at an AI security firm found the bug, and when the firm's autonomous system raced its own people to write the exploit, the people won again, with the machine getting there only against a deliberately weakened target. That is precisely the model behind an effective hybrid program: machine breadth, human judgement at both ends. Note the other half of the researcher's own read, though, because it is the part that should shape your planning: he credited AI tools with helping humans understand unfamiliar code and investigate suspicious areas far faster. The window is shorter, not closed. For the evidence on where autonomous tools land today, see our AI pentest benchmark results and our note on the acceptable false-positive rate for autonomous pentesting.

Stingrai is a CREST-accredited penetration testing service provider, based in Canada, that tracks the AI-attacker era as a working reality rather than a forecast. Our hybrid model pairs autonomous coverage with senior human validation, and the resulting reports give you the pentest evidence that supports your SOC 2 and ISO 27001 programs. Scope and packages are on our pricing page.

Frequently Asked Questions

Which real CVEs did AI discover in 2026?

Four verified records carry an AI finder credit across 2025 and 2026: the Microsoft Devices Pricing Program RCE CVE-2026-21536 (XBOW, autonomous), the SQLite flaw CVE-2025-6965 (Google Big Sleep, autonomous), the FFmpeg cluster CVE-2026-39210 through CVE-2026-39218 (depthfirst, autonomous), and the Apple macOS flaw CVE-2026-28952 (Calif.io with Claude and Anthropic Research, AI-assisted). A fifth record, the Exim RCE CVE-2026-45185, is tracked because it is routinely miscredited to AI: it was discovered and reported by a human, XBOW Security Lab head Federico Kirschbaum, and the AI only raced on the exploit that followed, where the human won and the autonomous system succeeded only against a deliberately weakened target. All five are patched.

What is the highest-severity CVE credited to an AI in 2026?

CVE-2026-21536, a remote code execution flaw in the Microsoft Devices Pricing Program, carries a CVSS 3.1 base score of 9.8 and was credited to the autonomous agent XBOW (NVD). The Exim RCE CVE-2026-45185 also scores 9.8, but its finder credit is human, so it is not an AI find at all.

Did an AI agent really find a bug that had existed for 23 years?

Yes. Among the 21 FFmpeg zero-days that depthfirst's agent reported in June 2026, one stack overflow in service-description-table handling dated to 2003 and had gone undisclosed for roughly 23 years (The Hacker News, 2026). It is a clear example of agents finding value in mature code where manual review had moved on.

What is the difference between an autonomous AI find and an AI-assisted find?

An autonomous find means an agent drove the discovery end to end, with a human only triaging before disclosure, as with XBOW on Microsoft or Big Sleep on SQLite. An AI-assisted find means a human researcher used a model as a tool and is credited alongside it, as with the Apple macOS entry. The third bucket, human-found and AI-raced, is where a person made the discovery and an AI system only competed on the follow-on exploit work, as with the Exim RCE. The three are different threat claims and should not be merged.

Why is the Kimi K3 and Redis finding not in the tracker?

Because the attribution does not meet the bar. The separate claim that Kimi K3 found a Redis issue quickly has no CVE assigned to it, and the numbered Redis CVEs that circulate with that story carry named human credits, with the one that has a genuine AI angle, CVE-2026-23479, belonging in the AI-assisted bucket rather than the autonomous one, because Team Xint Code found it by running Theori's autonomous code auditing tool Xint Code. The tracker only records numbered, fixed vulnerabilities with a verifiable finder, so this pairing does not qualify yet.

Are all of these AI-attributed CVEs patched?

Yes. Every entry points to a fixed version or a server-side mitigation: Microsoft mitigated CVE-2026-21536 in its cloud, Exim fixed CVE-2026-45185 in 4.99.3, SQLite fixed CVE-2025-6965 in 3.50.2, the FFmpeg cluster is fixed upstream, and Apple shipped fixes for CVE-2026-28952 across macOS, iOS, and iPadOS. This is a patch-status reference, not an open-threat feed.

Does an AI-found CVE mean AI can now replace human pentesters?

No. The 2026 record shows autonomous agents reaching real severity, but on the Exim RCE a human found the bug and humans then beat the autonomous system to a working exploit, inside XBOW, the AI security firm that ran both sides, and the machine landed an exploit only against a deliberately weakened target. Federico Kirschbaum, who heads XBOW's Security Lab and found the Exim bug, said he does not think "LLMs alone are quite ready to write exploits against real-world software" yet, and in the same account he credited AI tools with a crucial role in helping humans understand unfamiliar code and investigate suspicious areas far faster (BleepingComputer, 2026; The Hacker News, 2026). The effective model is hybrid: autonomous breadth plus senior human validation, which is how Stingrai runs its testing.

Where can I get the latest AI-attributed CVE data?

Track the primary advisories directly: NVD for per-CVE detail, Apple's own security-content advisories, and lab and vendor write-ups from Google Big Sleep, XBOW, and depthfirst. Stingrai's AI-attributed CVE tracker is refreshed as new dated, verifiable AI-attributed CVEs publish, and every row links back to its primary advisory so any claim can be audited inline.

References

  1. Apple. About the security content of macOS Tahoe 26.5. May 2026. https://support.apple.com/en-us/127115. Vendor advisory listing CVE-2026-28952 and its credit to Calif.io in collaboration with Claude and Anthropic Research.

  2. NVD (NIST). CVE-2026-28952. 2026. https://nvd.nist.gov/vuln/detail/CVE-2026-28952. Integer overflow, CVSS 3.1 base 7.5, fixed across macOS, iOS, and iPadOS builds.

  3. XBOW. Three RCE vulnerabilities in Microsoft identified by XBOW. March 2026. https://xbow.com/blog/three-rce-vulnerabilities-in-microsoft-identified-xbow. Vendor write-up on the autonomous discovery of the Microsoft RCE cluster.

  4. Krebs on Security. Microsoft Patch Tuesday, March 2026 Edition. March 2026. https://krebsonsecurity.com/2026/03/microsoft-patch-tuesday-march-2026-edition/. Independent coverage of the March 2026 bulletin, including CVE-2026-21536.

  5. NVD (NIST). CVE-2026-21536. 2026. https://nvd.nist.gov/vuln/detail/CVE-2026-21536. Unrestricted file upload (CWE-434), CVSS 3.1 base 9.8, Microsoft Devices Pricing Program.

  6. Google Cloud. Cloud CISO Perspectives: our Big Sleep agent makes a big leap. 2025. https://cloud.google.com/blog/products/identity-security/cloud-ciso-perspectives-our-big-sleep-agent-makes-big-leap. Google's account of Big Sleep finding SQLite CVE-2025-6965 and foiling an in-the-wild exploit.

  7. NVD (NIST). CVE-2025-6965. 2025. https://nvd.nist.gov/vuln/detail/CVE-2025-6965. SQLite memory corruption before 3.50.2, CVSS 3.1 base 7.7 and CVSS 4.0 base 7.2.

  8. depthfirst. 21 Zero-Days in FFmpeg. June 2026. https://depthfirst.com/research/21-zero-days-in-ffmpeg. Research write-up on an autonomous agent finding 21 FFmpeg zero-days for roughly US$1,000.

  9. The Hacker News. AI Agent Uncovers 21 Zero-Days in FFmpeg; Chrome Patches Record 429 Bugs. June 2026. https://thehackernews.com/2026/06/ai-agent-uncovers-21-zero-days-in.html. Coverage confirming the CVE-2026-39210 through -39218 range and the 2003-era stack overflow.

  10. NVD (NIST). CVE-2026-39210. 2026. https://nvd.nist.gov/vuln/detail/CVE-2026-39210. Reserved entry heading the FFmpeg cluster; per-CVE severities pending NVD analysis.

  11. BleepingComputer. New critical Exim mailer flaw allows remote code execution. May 2026. https://www.bleepingcomputer.com/news/security/new-critical-exim-mailer-flaw-allows-remote-code-execution/. Independent coverage crediting Federico Kirschbaum, head of Security Lab at XBOW, with discovering and reporting CVE-2026-45185 on 1 May 2026, and the source for his personal assessment of what LLMs can and cannot yet do on exploit writing.

  12. The Hacker News. New Exim BDAT Vulnerability Exposes GnuTLS Builds to Potential Code Execution. May 2026. https://thehackernews.com/2026/05/new-exim-bdat-vulnerability-exposes.html. Independent coverage of CVE-2026-45185, the human finder credit, and the exploit-writing race that followed.

  13. NVD (NIST). CVE-2026-45185. 2026. https://nvd.nist.gov/vuln/detail/CVE-2026-45185. Exim use-after-free (CWE-416) in the BDAT path on GnuTLS builds, CVSS 3.1 base 9.8, fixed in 4.99.3.

  14. XBOW. Dead.Letter: how XBOW found an unauthenticated RCE on Exim. May 2026. https://xbow.com/blog/dead-letter-cve-2026-45185-xbow-found-rce-exim. Vendor write-up and origin of the account, covering the internal exploit-writing race and the autonomous system's exploit against a simplified target. Listed after the trade reports because xbow.com returned HTTP 429 to every automated request on our verification passes, so it is not currently reachable for independent checking.

0 views

0

X

Related reading

Who Found the Bug? Why the CVE Record Cannot Say 'An AI Did It', and What That Breaks
AdvisoriesLLM Security

Who Found the Bug? Why the CVE Record Cannot Say 'An AI Did It', and What That Breaks

The CVE credits field records people, organizations and tools, never autonomy. How to verify an AI-discovered vulnerability claim against the real record.

11 min read

Nine CVEs in the Tools Your Developers Run All Day: Cursor, Claude Code, Gemini CLI and the MCP Reference Server
AdvisoriesLLM Security

Nine CVEs in the Tools Your Developers Run All Day: Cursor, Claude Code, Gemini CLI and the MCP Reference Server

Nine 2026 AI coding assistant CVEs across Cursor, Claude Code, Gemini CLI and mcp-server-git: version floors, CI hardening and what to test now.

12 min read

Jailbroken, Misused, Flawed or Breached: A Category-by-Category Reading of the 2026 AI Company Hacked Headlines
LLM SecurityAdvisories

Jailbroken, Misused, Flawed or Breached: A Category-by-Category Reading of the 2026 AI Company Hacked Headlines

Jailbreak, misuse, product bug or real breach? A category test of the 2026 AI vendor incidents shows which ones put your data at risk and which do not.

17 min read

Contents

X