main logo icon

Published on

July 23, 2026

|

13 min read

Anthropic Mapped a Year of AI Attacks to MITRE ATT&CK and Found the Layer It Is Missing

A defender's breakdown of Anthropic's year-long dataset of AI-enabled cyber threats mapped to MITRE ATT&CK: 832 accounts, 13,873 actions, 482 techniques, and the agentic-orchestration layer the framework has not yet codified.

Arafat Afzalzada

Arafat Afzalzada

Founder

LLM SecurityNetwork Security

Summarize with AI

ChatGPTPerplexityGeminiGrokClaude

TL;DR

- The dataset (Mar 2025 to Mar 2026): Anthropic analyzed 832 accounts it banned for malicious cyber activity, logging 13,873 actions across 482 unique MITRE ATT&CK techniques and all 14 tactics (Anthropic, 2026). - The scoped answer: AI-enabled attacks are not outrunning ATT&CK as a taxonomy. Every observed action mapped into the existing 14 tactics. What ATT&CK has not yet codified is the agentic-orchestration layer, the way an autonomous agent chains those techniques on its own (Anthropic, 2026). - Risk is climbing fast: medium-risk-or-higher actors rose from 33% in the first six months to 56% in the second, a roughly 1.7x increase (Anthropic, 2026). - Use of AI moved deeper into the kill chain: AI-assisted account discovery inside compromised environments rose 8.9% while AI-assisted phishing fell 8.6% (Anthropic, 2026). - Capability development dominated: 67.3% of the 832 accounts (560) used AI to develop malware or otherwise prepare an attack (Anthropic, 2026). - The techniques attackers leaned on: Develop Capabilities (T1587) was the most common family at 69% of actors, defense evasion was the largest tactic category at 84.4%, and only 6.5% used AI for lateral movement (Anthropic, 2026). - The missing layer, named: autonomous killchain orchestration, real-time pivot decisions, and AI-directed execution with no human intervention have no ATT&CK IDs yet (Anthropic, 2026). - What defenders should do: stop grading coverage on technique IDs alone, add orchestration-layer signals to threat models and purple-team plans, and test for the agentic behaviors the matrix cannot yet name.

Anthropic analyzed 832 accounts it banned for malicious cyber activity between March 2025 and March 2026, mapped every observed behavior to the MITRE ATT&CK framework, and logged 13,873 actions across 482 unique techniques and all 14 ATT&CK tactics (Anthropic). Read one way, that is reassuring: a full year of real AI-enabled attacks fit inside the taxonomy defenders already use. Read another way, it exposes a gap. The behaviors that separated the most dangerous actors from the rest were not new techniques. They were the way an autonomous agent chained existing techniques together on its own, and that orchestration layer has no ATT&CK ID. As Anthropic put it, "there is no ATT&CK ID for this type of agentic orchestration" (Anthropic).

So here is the direct answer to the question every detection team is asking. AI-enabled cyberattacks are not outrunning MITRE ATT&CK as a classification of individual techniques. What they are outrunning is the framework's ability to describe how autonomous agents sequence those techniques. Three forces make that gap urgent in 2026. Medium-risk-or-higher actors jumped from 33% to 56% across the study, a roughly 1.7x increase (Anthropic). Attacker use of AI moved deeper into the kill chain, with account discovery up 8.9% and AI-assisted phishing down 8.6% (Anthropic). And the highest-risk operation Anthropic disrupted stood out for its orchestration, not its technique count (Help Net Security). For CISOs, detection engineers, and red-team buyers, that reframes the coverage question entirely.

This post is the Stingrai research team's canonical 2026 breakdown of Anthropic's ATT&CK-mapped dataset. It carries 12 attributed figures drawn from three primary publishers: Anthropic's first-party threat-intelligence writeup, its companion ATT&CK Navigator research page, and corroboration from Help Net Security. The lead data covers Anthropic's study window of March 2025 to March 2026, published June 2026, which is the freshest first-party analysis of its kind available as of July 2026. Every figure below carries its source, its year, and the exact behavior it describes so any claim can be audited inline. This is a data breakdown, not a coverage-grading how-to; for the buyer's rubric on scoring ATT&CK coverage in a proposal, see the companion piece linked near the end.

TL;DR: the dataset by the numbers

  • The corpus (Mar 2025 to Mar 2026): 832 accounts banned for malicious cyber activity, 13,873 logged actions, 482 unique ATT&CK techniques, all 14 tactics (Anthropic, ATT&CK Navigator).

  • The scoped finding: every action mapped into ATT&CK's existing 14 tactics. The uncodified part is agentic orchestration, not the techniques themselves (Anthropic).

  • Risk climbed 1.7x: medium-risk-or-higher actors rose from 33% in the first six months to 56% in the second (Anthropic).

  • AI moved post-compromise: account discovery inside compromised environments rose 8.9%, AI-assisted phishing fell 8.6% (Anthropic).

  • Capability development led: 67.3% of accounts (560 of 832) used AI to develop malware or prepare an attack (Anthropic).

  • Most common technique family: Develop Capabilities (T1587), used by 574 of 832 actors, or 69% (Anthropic, ATT&CK Navigator).

  • Largest tactic category: defense evasion, present in the behavior of 84.4% of the actors studied (Anthropic, ATT&CK Navigator).

  • Lateral movement stayed rare: only 54 of 832 actors (6.5%) used AI to assist lateral movement (Anthropic, ATT&CK Navigator).

  • The named gap: autonomous killchain orchestration, real-time pivot decisions, and AI-directed execution with no human intervention have no ATT&CK ID numbers yet (Anthropic, ATT&CK Navigator).

Key takeaways

AI attacks fit the ATT&CK matrix; their orchestration does not. The single most misread headline of the summer is that AI has broken ATT&CK. It has not. A full year of banned activity mapped cleanly into 482 existing techniques and all 14 tactics (Anthropic). The gap is narrow and specific: the framework enumerates what an attacker does, not how an autonomous agent decides, in real time and without a human, to chain those actions together.

Risk is scaling faster than skill. The share of medium-risk-or-higher actors rose from 33% to 56% in a single year (Anthropic). AI is lowering the floor, letting less capable actors reach further into techniques that once demanded expertise, which means your threat model can no longer assume that an advanced technique implies an advanced adversary.

The action moved to where you have the least coverage. AI use shifted from initial access toward post-compromise work: account discovery inside compromised environments rose while phishing fell (Anthropic). Many detection programs are still weighted toward the perimeter. The data says the machine-assisted activity is increasingly happening after the breach.

Technique count is a bad proxy for danger. The highest-risk operation Anthropic studied was distinct "not because of the number of techniques it employed but because of how the attackers used an AI agent to orchestrate them" (Help Net Security). A coverage matrix that grades on how many technique IDs a test touches will systematically miss the thing that actually made that actor dangerous.

Methodology and sourcing

The figures in this post come from three sources, all cited inline. The primary dataset is Anthropic's first-party analysis of 832 accounts it banned for violating the cyber-related parts of its usage policy between March 2025 and March 2026, published as "What we learned mapping a year's worth of AI-enabled cyber threats" on June 3, 2026, and its companion research page, the LLM ATT&CK Navigator. Corroboration comes from Help Net Security's June 5, 2026 writeup by managing editor Sinisa Markovic. Where the two Anthropic pages report a figure at different granularity, this post keeps them separate: the 67.3% "malware development" figure describes an activity, while the 69% Develop Capabilities (T1587) figure describes a mapped technique family, and the two are labeled distinctly rather than merged. Every source URL returned a live page during the research pass on July 23, 2026. Figures that could not be traced to one of these named primary sources were left out rather than estimated, and no full-year 2026 numbers are claimed because none exist yet for this study.

Are AI-enabled cyberattacks outpacing MITRE ATT&CK?

Not as a taxonomy of techniques. This is the claim to get right, because the sensational version travels faster than the accurate one. Anthropic observed 13,873 discrete actions and every one of them mapped to an existing ATT&CK technique inside the existing 14 tactics (Anthropic). If AI attacks were escaping the framework, you would expect a long tail of behaviors that had no home in the matrix. That is not what the data shows. The matrix held.

What the matrix does not hold is the orchestration. ATT&CK is a catalog of discrete techniques and the tactics they serve. It was built to answer "what did the adversary do," and it answers that well even for AI-assisted operations. It was not built to answer "who, or what, decided the order, and how fast, and with how little human input." Anthropic names the specific behaviors that fall through: "autonomous killchain orchestration, real-time pivot decisions, and AI-directed execution with no human intervention don't yet have ID numbers in the ATT&CK framework" (Anthropic). Those are not new techniques. They are a property of how the techniques are strung together.

For a detection team, the practical consequence is that a coverage matrix keyed purely on technique IDs can show green across the board and still be blind to the thing that makes an agentic adversary fast and scalable. Your threat model needs a second axis: not just which techniques you can detect, but whether you can detect them being chained autonomously, at machine tempo, without a human in the loop.

What the dataset actually shows

Two shifts inside the year matter more than the totals. The first is risk concentration. In the first six-month period, 33% of actors were classified as medium risk or higher by Anthropic's risk-scoring system. By the second six-month period, that share had jumped to 56%, a roughly 1.7x increase (Anthropic). That is not a story about a handful of elite crews getting better. It is a story about the middle of the distribution moving up, as AI lets less-skilled actors operate at a level that used to require real expertise.

The second shift is where in the kill chain the AI got used. Across the period, attacker use of AI "shifted from techniques to gain initial access to a system towards activity carried out once they were inside the system" (Anthropic). The clearest signal is a pair of opposing moves: AI-assisted account discovery, identifying valid accounts inside a compromised environment, rose 8.9%, while AI-assisted phishing, a classic way in, fell 8.6% (Anthropic).

Anthropic Attck Risk Shift

For detection engineering, that pivot is the actionable part. If AI-assisted activity is migrating from the front door to the interior, then coverage that is heavily weighted toward inbound phishing and initial-access controls is defending the phase attackers are leaning on less. The relative growth is in discovery, credential use, and quiet movement inside an environment that already trusts the session. Those are exactly the behaviors that hide in normal administrative noise, and exactly where post-compromise detection and identity telemetry earn their keep.

The tactics and techniques attackers leaned on

Broken down by observed behavior, the 832 accounts show a lopsided profile. The most common technique family was Develop Capabilities (T1587), used by 574 of the 832 actors, or 69% (Anthropic). Closely related, 67.3% of accounts (560 of 832) used AI to develop malware or otherwise prepare an attack (Anthropic). The single largest tactic category was defense evasion, present in the behavior of 84.4% of the actors studied (Anthropic). And at the other end, only 54 of 832 actors, or 6.5%, used AI to assist lateral movement (Anthropic).

Anthropic Attck Tactics

The shape tells a defender two things. First, attackers are overwhelmingly using AI for the preparation and evasion work: building tooling, developing malware, and staying quiet. That is the labor-intensive part of an operation, and it is precisely the part a model is good at accelerating. Second, the low lateral-movement number is a snapshot, not a ceiling. Anthropic frames these autonomous, chained behaviors as "precisely the behaviors we expect to see much more of as AI agents become more capable" (Anthropic). The 6.5% is where the frontier sits today, and it is the number most likely to move next.

Note the granularity difference, because it matters for honest reading: defense evasion is a tactic, Develop Capabilities is a technique family, and malware development is an activity. They are all shares of the same 832 accounts, which is why they sit on one chart, but they describe different levels of the ATT&CK hierarchy. This is the kind of distinction a rigorous coverage claim preserves and a padded one blurs.

The layer ATT&CK is missing: agentic orchestration

The clearest illustration of the gap is the highest-risk actor in the dataset. Anthropic's analysis found that technique count or tactic type alone could not explain what made that operation, tracked as GTG-1002, so dangerous. What set it apart was orchestration: the actor "weaponized Claude Code running on a Kali Linux machine, integrating open-source penetration testing tools as MCP (Model Context Protocol) servers" so the AI could drive the tooling and chain steps with minimal human direction (Anthropic). Help Net Security summarizes the point crisply: the attack "was distinct not because of the number of techniques it employed but because of how the attackers used an AI agent to orchestrate them" (Help Net Security).

Every individual technique in that operation has an ATT&CK ID. The orchestration does not. That is the missing layer, and it is worth stating its boundaries precisely so nobody overclaims. ATT&CK is not obsolete, and AI attacks are not off the map. The framework simply catalogs techniques, and the new variable is a control layer sitting above the techniques: an agent that plans, executes, observes the result, and re-plans, at a tempo and autonomy level the ID scheme was never designed to express. Stingrai's deeper look at the GTG-1002 operation, from a defender's point of view, is worth reading alongside this breakdown; it is linked in the references and in the closing section.

Which agentic techniques your coverage matrix misses

A coverage matrix built as a spreadsheet of detectable technique IDs structurally cannot represent the following, drawn straight from the behaviors Anthropic names as uncodified (Anthropic):

  • Autonomous killchain orchestration. A single agent advancing from discovery to execution across multiple tactics without a human choosing each next step. Your matrix scores each technique in isolation; it has no field for "these ran as one autonomous chain."

  • Real-time pivot decisions. The agent observing a failed action or a new opening and re-planning mid-operation. This is a property of the decision loop, not a technique, so no single ID captures it.

  • AI-directed execution with no human intervention. Tools invoked by the model itself, for example through MCP servers, rather than by an operator at a keyboard. The technique that runs is the same; who ran it, and how fast, is what changed.

  • Machine-tempo sequencing. Dozens of techniques executed faster than a human team could, compressing the dwell time your playbooks assume. Tempo is not a technique either.

None of these replace ATT&CK. They ride on top of it. The defensive move is to add an orchestration axis to your threat model and your purple-team plan: for each critical attack path, ask not only "can we detect technique X" but "can we detect X, Y, and Z being chained autonomously and at speed, and would our response fire before the agent finishes the sequence." That is the question the coverage matrix, on its own, cannot answer.

What this means for defenders

  • Stop treating technique-ID coverage as the whole score. A matrix that is all green on individual techniques can still miss an agentic adversary entirely, because the danger was in the orchestration, not the technique list (Help Net Security). Grade depth and sequencing, not just breadth. Our companion rubric on reading ATT&CK coverage in a red team proposal walks through exactly how to spot padding.

  • Rebalance detection toward post-compromise. With AI-assisted discovery rising and phishing falling (Anthropic), invest in identity telemetry, account-discovery detections, and internal movement analytics, not only perimeter and email controls. Benchmark what your blue team actually catches with red-team detection benchmarks.

  • Add an orchestration axis to threat models and purple teams. For each crown-jewel path, test whether you can detect techniques being chained autonomously and at machine tempo, not just in isolation. Choosing which adversary to model first is easier with a threat-group-by-industry guide.

  • Assume advanced techniques no longer imply advanced adversaries. With medium-risk-or-higher actors up 1.7x in a year (Anthropic), your model should expect capable-looking behavior from a broader population. The wider AI cyberattack statistics for 2026 put that trend in context.

  • Test continuously, because the frontier moves inside a single year. The 6.5% lateral-movement figure is the one Anthropic expects to grow (Anthropic). Point-in-time snapshots go stale between the shifts. This is the case for continuous red teaming over the annual pentest.

Where Stingrai fits

Stingrai runs threat-led, ATT&CK-mapped red teaming that is built by senior human operators, and that human-led design is exactly what lets an engagement model the orchestration layer the framework has not yet codified. Rather than checking off technique IDs, Stingrai's red team ties each technique to a named objective and, where the threat model calls for it, emulates the agentic patterns Anthropic flags: autonomous chaining across tactics, real-time re-planning, and AI-directed tooling driven with minimal human input. The deliverable is not a padded matrix but evidence of what your blue team detected, and what it missed, as those sequences ran.

That maps to a continuous posture rather than a once-a-year snapshot. Because the data shows the frontier moving inside a single year, Stingrai's continuous red-team model re-tests as the tradecraft shifts, and the resulting ATT&CK coverage benchmarking supports the evidence your SOC 2, ISO 27001, and DORA programs need to show that controls were exercised against current adversary behavior. Stingrai has been building offensive-security capability since 2021 as a CREST-accredited penetration testing service provider, from its Toronto headquarters and London office. For the web-application surface specifically, Stingrai's autonomous agent, Snipe, hunts complex classes like IDOR, business-logic, and broken-authorization flaws; the human-led red team is what covers the agentic-orchestration behaviors this post is about. You can see how the engagements are packaged on the Stingrai pricing page and across its offensive security services.

Frequently asked questions

Are AI-enabled cyberattacks outpacing MITRE ATT&CK, and which agentic techniques does my coverage matrix miss?

Not as a taxonomy. Anthropic mapped a full year of AI-enabled attacks, 13,873 actions from 832 banned accounts, into ATT&CK's existing 482 techniques and all 14 tactics, so the techniques themselves fit the framework (Anthropic, 2026). What ATT&CK has not yet codified is the agentic-orchestration layer: autonomous killchain orchestration, real-time pivot decisions, and AI-directed execution with no human intervention have no ATT&CK IDs. A coverage matrix built on technique IDs will miss those because they are properties of how techniques are chained, not techniques themselves.

How many accounts and actions did Anthropic analyze?

Anthropic analyzed 832 accounts it banned for malicious cyber activity between March 2025 and March 2026, logging 13,873 actions across 482 unique techniques and all 14 ATT&CK tactics (Anthropic, 2026). It is the first year-scale, first-party mapping of AI-enabled attacker behavior to ATT&CK at this volume.

Did AI-enabled attacker risk actually increase over the year?

Yes. The share of actors classified as medium risk or higher rose from 33% in the first six-month period to 56% in the second, a roughly 1.7x increase (Anthropic, 2026). The trend reflects AI lowering the barrier so that less capable actors can operate at a higher level, not just elite groups improving.

What did attackers use AI for most?

Capability development. 67.3% of the 832 accounts (560) used AI to develop malware or otherwise prepare an attack, and Develop Capabilities (T1587) was the most common technique family at 69% of actors (Anthropic, 2026). Defense evasion was the largest tactic category, present in 84.4% of actors.

Where in the attack lifecycle is AI use growing?

Post-compromise. Anthropic found AI use shifting from initial access toward activity carried out once attackers were already inside a system. AI-assisted account discovery inside compromised environments rose 8.9%, while AI-assisted phishing fell 8.6% (Anthropic, 2026). For defenders, that argues for weighting detection toward identity and internal-movement telemetry.

What is agentic orchestration, and why does it lack an ATT&CK ID?

Agentic orchestration is an autonomous agent planning, executing, observing results, and re-planning across multiple ATT&CK techniques on its own, at machine tempo, with little or no human direction. ATT&CK catalogs discrete techniques and the tactics they serve; it does not have IDs for the control layer that sequences them autonomously. Anthropic names three such behaviors that have no ATT&CK IDs yet: autonomous killchain orchestration, real-time pivot decisions, and AI-directed execution with no human intervention (Anthropic, 2026).

Does this mean MITRE ATT&CK is obsolete?

No. Every one of the 13,873 observed actions mapped to an existing ATT&CK technique, so the framework remains the right backbone for detection coverage (Anthropic, 2026). The gap is narrow: ATT&CK describes techniques well and does not yet describe how autonomous agents chain them. Defenders should keep using ATT&CK and add an orchestration axis on top of it, not replace it.

How does a red team test for behaviors that have no ATT&CK ID?

By emulating the orchestration, not just the techniques. A threat-led red team can chain techniques autonomously and at speed against a defined objective, then report what the blue team detected and missed as the sequence ran. That is human-led adversary emulation designed around the objective and the tempo, which is how Stingrai's red-team engagements exercise the agentic patterns the matrix cannot yet name.

Where can I read the primary data myself?

Anthropic published both a narrative writeup, "What we learned mapping a year's worth of AI-enabled cyber threats," and a companion research page, the LLM ATT&CK Navigator, both linked in the references below. Help Net Security's June 5, 2026 article provides an independent summary.

References

  1. Anthropic. What we learned mapping a year's worth of AI-enabled cyber threats. June 3, 2026. https://www.anthropic.com/news/AI-enabled-cyber-threats-mitre-attack. First-party analysis of 832 banned accounts; source of the 33% to 56% risk shift, the 67.3% malware-development figure, the account-discovery and phishing deltas, and the "no ATT&CK ID for this type of agentic orchestration" statement.

  2. Anthropic. LLM ATT&CK Navigator (companion research page). 2026. https://www.anthropic.com/research/attack-navigator. Source of the 13,873 actions / 482 techniques / 14 tactics total, the T1587 (69%) and defense-evasion (84.4%) breakdowns, the 6.5% lateral-movement figure, the GTG-1002 orchestration detail, and the three uncodified agentic behaviors.

  3. Help Net Security (Sinisa Markovic). AI is helping low-skill hackers pull off advanced cyberattacks. June 5, 2026. https://www.helpnetsecurity.com/2026/06/05/anthropic-ai-cyber-activity-analysis/. Independent trade-press corroboration of the 832-account dataset, the totals, the risk shift, and the orchestration-over-technique-count framing.

  4. MITRE. ATT&CK for Enterprise. https://attack.mitre.org/. The adversary tactics-and-techniques knowledge base every figure in this post is mapped against.

  5. Stingrai. Anthropic, Mythos, and GTG-1002: a defender's analysis. https://www.stingrai.io/blog/anthropic-mythos-gtg1002-defender-analysis. Deeper defender-side reading on the GTG-1002 operation referenced above.

  6. Stingrai. How to Read ATT&CK Coverage in a Red Team Proposal. https://www.stingrai.io/blog/grading-mitre-attack-coverage-red-team-proposal. The buyer's rubric for grading ATT&CK coverage and spotting technique-count padding.

0 views

0

X

Related reading

PROMPTSTEAL and PROMPTFLUX: Malware That Calls an LLM to Attack
LLM SecurityNetwork Security

PROMPTSTEAL and PROMPTFLUX: Malware That Calls an LLM to Attack

PROMPTSTEAL and PROMPTFLUX are malware that query an LLM at runtime. See how APT28 weaponizes AI in live attacks and how blue teams detect LLM driven code.

15 min read

From Headline to Test Case: Mapping 2026's AI-Attacker Milestones to Red Team Coverage
LLM SecurityNetwork Security

From Headline to Test Case: Mapping 2026's AI-Attacker Milestones to Red Team Coverage

Map 2026's real AI attacker milestones to the red team objectives and concrete test cases they imply, from autonomous agent breaches to AI found zero days.

17 min read

AI Cyber Attack Statistics 2026: Attacker AI, Influence Ops, and Agentic Threats
LLM SecurityNetwork Security

AI Cyber Attack Statistics 2026: Attacker AI, Influence Ops, and Agentic Threats

AI Cyber Attack Statistics 2026. 1 in 6 breaches use AI. GTG-1002 ran 80-90 percent of attack ops. Verified data from Anthropic, IBM, Microsoft, and more.

24 min read

Contents

X