Article 15 of the EU AI Act requires every high-risk AI system to be resilient against data poisoning, model poisoning, adversarial examples (also called model evasion), and confidentiality attacks or model flaws, and to resist unauthorised third parties who try to alter its use, outputs, or performance by exploiting system vulnerabilities (EU AI Act, Article 15(5)). The Act never uses the phrase penetration testing. It does name the attack classes your system must withstand, and the practical way a provider or deployer shows it addressed them is security testing evidence, filed in the system's technical documentation.
Does the EU AI Act require penetration testing of high-risk AI systems, and what evidence should you have ready before the December 2027 deadline? The short answer: the law does not mandate a test called a penetration test, but it does mandate the security properties that adversarial robustness testing and application penetration testing produce evidence for. To satisfy Article 15 you need to be able to show, in writing, that you assessed your system against each named attack class and mitigated what you found. That is procurement, not paperwork you can generate at a desk, and the window to procure it is open now.
As of July 2026: the Digital Omnibus that moved these deadlines was adopted by the European Parliament on 16 June 2026 and approved by the Council of the EU on 29 June 2026, and the final act was signed on 8 July 2026. It is pending publication in the Official Journal and enters into force on the third day after publication. The new dates below bind once that publication happens; treat them as firm and build against them.
TL;DR
Attack classes named in law (Article 15(5)): data poisoning, model poisoning, adversarial examples / model evasion, and confidentiality attacks or model flaws (EU AI Act, Article 15).
New application date, standalone Annex III high-risk systems: 2 December 2027, deferred from the original 2 August 2026 (Gibson Dunn analysis of the Omnibus).
New application date, embedded Annex I product high-risk systems: 2 August 2028, deferred from 2 August 2027 (Gibson Dunn).
Parliament adopted the Digital Omnibus: 16 June 2026, by 423 votes to 57 with 174 abstentions (Council of the EU).
Council gave final approval: 29 June 2026 (Council of the EU).
Status as of July 2026: signed 8 July 2026, pending Official Journal publication; the deferred dates take legal effect once published (EU AI Act news roundup).
What did not change: the substance of Article 15. The Omnibus deferred timelines only (Freshfields).
The evidence that satisfies it: poisoning-resilience testing, adversarial robustness testing, confidentiality-attack testing, and application and API penetration testing of the AI system, each producing a report you file.
Why it matters now: 13% of organisations reported a breach of their AI models or applications in the past year, and 97% of those lacked proper AI access controls, per IBM's 2025 Cost of a Data Breach Report.
Key takeaways
The deadline moved, the obligation did not. The most common misread of the 2026 Omnibus is that pressure came off. It did not. Standalone Annex III systems apply from 2 December 2027 and embedded product systems from 2 August 2028, but the Article 15 requirements they must meet are unchanged. The extra runway is procurement time, not a reason to wait.
Article 15 is a design-and-evidence duty, not a certificate. No regulator issues an Article 15 pass. You demonstrate conformity through your own technical documentation, and security testing reports are the strongest evidence in that file that you addressed the named attack classes.
The four named attack classes map cleanly to four procurable tests. Data poisoning, model poisoning, model evasion, and confidentiality attacks each correspond to a specific test type with a specific deliverable. The matrix below turns the clause into a shopping list.
Adversarial robustness testing and application penetration testing are complementary, not interchangeable. One probes the model's behaviour under crafted input; the other probes the application, API, and access controls wrapped around it. Article 15(5) names both surfaces, so your evidence pack needs both.
Providers and deployers carry different weight. The provider builds and documents the system, so it owns most of the Article 15 evidence. A deployer that materially modifies a high-risk system, or fine-tunes a model on its own data, can inherit provider-level duties and should procure its own testing.
How to read this guide
This is the Stingrai security team's working reference for the security testing side of Article 15, written for providers and deployers of high-risk AI systems who need to know what to procure and file during the grace window. Every legal claim links to the Article 15 text or to the official Parliament and Council record so any statement can be checked at the source. Dates reflect the Digital Omnibus as adopted and signed in mid-2026; because Official Journal publication was still pending when this was written, we date the status plainly rather than imply the deferral is already in force. Nothing here is legal advice; it is a procurement map for the testing evidence Article 15 makes you responsible for.
Does the EU AI Act require penetration testing of high-risk AI systems?
The EU AI Act does not name penetration testing as a mandatory activity. Article 15 sets an outcome: a high-risk AI system must achieve an appropriate level of accuracy, robustness, and cybersecurity, and perform consistently across those dimensions throughout its lifecycle. Article 15(5) then makes the cybersecurity outcome concrete by naming the attack classes the system must be resilient to. It requires, "where appropriate," technical measures to prevent, detect, respond to, resolve, and control for those attacks.
That phrasing is why the honest answer is nuanced. You are not obliged to run a test with a particular label. You are obliged to build a system that resists the named attacks and to be able to evidence that you did. In practice, the way an organisation shows it prevented and detected data poisoning, evasion, and confidentiality attacks is a test that attempts exactly those attacks and a report that records the result. That is penetration testing and adversarial robustness testing by function, whatever the invoice calls it. Regulators, auditors, and enterprise customers reading your technical documentation will look for that evidence, and its absence is the gap that turns an audit into a finding.
What Article 15 actually says
Article 15 has three parts that matter for testing. Quoting the operative language keeps the mapping honest.
Accuracy, robustness, and cybersecurity as a lifecycle duty. Article 15(1): "High-risk AI systems shall be designed and developed in such a way that they achieve an appropriate level of accuracy, robustness, and cybersecurity, and that they perform consistently in those respects throughout their lifecycle." The phrase "throughout their lifecycle" is what makes a single pre-launch test insufficient on its own; models and their inputs drift.
Robustness and fail-safe design. Article 15(4) requires systems to be "as resilient as possible regarding errors, faults or inconsistencies," allows robustness to be achieved "through technical redundancy solutions, which may include backup or fail-safe plans," and requires systems that keep learning after deployment to address "feedback loops" with mitigation measures. This is the reliability half of the article.
Cybersecurity and the named attack classes. Article 15(5): "High-risk AI systems shall be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities." It then requires that technical solutions to address AI-specific vulnerabilities "shall include, where appropriate, measures to prevent, detect, respond to, resolve and control for attacks trying to manipulate the training data set (data poisoning), or pre-trained components used in training (model poisoning), inputs designed to cause the AI model to make a mistake (adversarial examples or model evasion), confidentiality attacks or model flaws."
Read the chapeau and the list together and you get two testing surfaces. The chapeau is classic security: unauthorised third parties exploiting vulnerabilities, which points at the application, API, identity, and access-control layer around the model. The list is AI-specific: attacks on the training data, the pre-trained components, the model's decision boundary, and the confidentiality of the model and its data. Your evidence pack needs to reach both.
The Article 15 clause-to-evidence matrix

This is the core of the guide. Each attack class named in Article 15(5), plus the robustness duty in 15(4), maps to a test you can buy and a deliverable you can file. Require every tester to state which Article 15 clause each finding maps to, the same way a mature AI pentest maps findings to the OWASP LLM and Agentic lists.
Attack class (Article 15) | What the law names | Test that produces the evidence | Deliverable for your evidence pack |
|---|---|---|---|
Data poisoning | Attacks that manipulate the training data set | Poisoning-resilience testing and a review of data-ingestion and provenance controls | Poisoning-resilience test report plus data-integrity control evidence |
Model poisoning | Attacks on pre-trained components used in training | Supply-chain review of pre-trained models, weights, and dependencies | Model and dependency provenance evidence with integrity checks |
Adversarial examples / model evasion | Inputs designed to cause the model to make a mistake | Adversarial robustness testing: crafted inputs that try to force a wrong or unsafe output | Adversarial robustness test report with attack success rates |
Confidentiality attacks | Attacks that extract the model or its data | Model inversion, membership inference, and model or data extraction testing | Confidentiality-attack test report |
Model flaws | Behavioural and logic flaws in the model itself | LLM red teaming: prompt injection, jailbreaks, and improper output handling | Red team findings report mapped to OWASP LLM risks |
Unauthorised third parties (15(5) chapeau) | Exploiting system vulnerabilities to alter use, outputs, or performance | Application and API penetration testing around the AI system: auth, access control, business logic | Penetration test report with severities and retest |
Robustness and fail-safe (15(4)) | Errors, faults, inconsistencies, and feedback loops | Robustness and resilience testing, including fail-safe and feedback-loop review | Robustness assessment and mitigation evidence |
The pattern to notice: the AI-specific classes are model-facing and best exercised by human-led AI red teaming, while the chapeau is application-facing and best exercised by penetration testing of the surface that wraps the model. A serious evidence pack has both, and the ultimate guide to adversarial inputs in LLMs goes deeper on the evasion and prompt-injection classes if you want the technical background before you scope.
Adversarial robustness testing versus penetration testing
The two most common test types in an Article 15 pack are easy to confuse and cover different clauses, so buyers should scope both explicitly.
Adversarial robustness testing targets the model's behaviour. It crafts inputs that try to flip a classification, evade a safety filter, extract training data, or leak the system prompt, and it measures how often those attempts succeed. This is the direct evidence for the adversarial examples, confidentiality attacks, and model flaws named in Article 15(5). It is a human-led discipline because the interesting attacks are creative and context-specific, and our AI red teaming guide for LLM and agentic apps covers how that engagement is run.
Penetration testing targets the system around the model: the web application, the API, authentication, authorisation, and the business logic that decides what a given user or agent is allowed to do. A model can be perfectly robust and the product still be breached because an IDOR flaw lets one tenant read another tenant's chat history, or because a broken authorisation check lets a low-privilege user invoke a high-privilege tool. That surface is exactly the "unauthorised third parties exploiting system vulnerabilities" language in the chapeau of 15(5), and it is tested the same way any application is tested. If your AI feature retrieves from a vector store, the access-control layer there is its own surface, covered in RAG and vector-store access-control testing.
Buy one without the other and your evidence pack has a hole a regulator or an enterprise customer can point to. Scope both, and require each to map its findings back to the Article 15 clauses.
The 2026 to 2028 preparation timeline

The grace window created by the Digital Omnibus is procurement time. Use it deliberately.
June 2026: the European Parliament adopted the Digital Omnibus on 16 June 2026 and the Council gave final approval on 29 June 2026 (Council of the EU).
July 2026: the act was signed on 8 July 2026 and is pending Official Journal publication; it enters into force on the third day after publication.
2026 to 2027, the grace window: scope and run the tests in the matrix, remediate what they find, and file the reports in your technical documentation. Because Article 15 is a lifecycle duty, plan for retesting, not a one-off.
2 December 2027: standalone Annex III high-risk systems, the use-case list that covers areas such as biometrics, critical infrastructure, education, employment, essential services, and credit, become subject to the full requirements (Gibson Dunn).
2 August 2028: high-risk AI embedded as a safety component in products already regulated under Annex I legislation becomes subject to the requirements.
The trap in a long runway is treating December 2027 as a date to start rather than a date to finish. Testing that surfaces a poisoning or authorisation flaw needs remediation and a retest before it becomes clean evidence, and that cycle takes months, not days. Working backward from December 2027, the sensible time to scope the first engagement is now.
What goes in your Article 15 evidence pack
The technical documentation for a high-risk system is where conformity is shown. For the Article 15 slice, a defensible pack contains:
A test report per named attack class. Poisoning resilience, adversarial robustness, confidentiality attacks, model-flaw red teaming, and application and API penetration testing, each mapped to the Article 15 clause it evidences.
Severity and remediation records. Each finding with a severity, a remediation, and a retest result that shows it was closed. Open criticals are the weakest thing an auditor can find in the file.
Coverage statements. A record of what was in scope: the model and version, the RAG and agent layers, the API, and the identity surface, so a reader can see nothing material was skipped.
A retest and monitoring plan. Because 15(1) makes this a lifecycle duty, the pack should show how testing repeats after significant changes. A retest trigger matrix is a clean way to define what change forces a fresh test.
A scope of work that a third party can read. The engagement scope itself is evidence of diligence. The AI and LLM pentest scope of work template gives you a fill-in-the-blank starting point that already maps to the OWASP LLM and Agentic lists.
Providers versus deployers: who owns the evidence
Article 15 obligations sit primarily with the provider, the entity that develops the high-risk system or has it developed and places it on the market under its own name. The provider builds the system, so it owns the bulk of the testing and documentation.
Deployers, the organisations that use a high-risk system in their operations, carry lighter Article 15 duties, but the line moves. A deployer that puts its name on the system, materially modifies it, or fine-tunes a model on its own data can take on provider-level responsibilities, and at that point it needs its own testing evidence rather than relying on the upstream vendor's. If you are procuring a high-risk AI product, ask the provider for its Article 15 testing evidence, and if you are customising that product, plan to test the parts you changed. The healthcare AI clinical LLM scoping guide works through this provider-versus-deployer split for one high-stakes sector.
How Stingrai supports your Article 15 evidence pack
Stingrai delivers the security testing that produces Article 15 evidence, framed as inputs to the documentation you own. An AI engagement with us is a hybrid. Senior human pentesters run the AI-specific, model-facing work: adversarial robustness testing, prompt injection and jailbreak red teaming, poisoning-resilience testing, and confidentiality-attack testing such as model inversion and membership inference. That is the direct evidence for the adversarial examples, confidentiality attacks, and model-flaw classes named in Article 15(5), delivered through our red teaming service.
Our autonomous web application agent, Snipe, covers the application and API layer that wraps the model: the surface Article 15(5) describes as unauthorised third parties exploiting system vulnerabilities. Snipe is built to hunt the complex classes generic scanners miss, IDOR, business logic flaws, and broken authorisation around the AI functionality, and it performs both black-box testing and white-box source review. You can see how that layer is delivered on our web application penetration testing and PTaaS pages.
Every finding is mapped to the Article 15 clause and the OWASP LLM or Agentic risk it evidences, with severity, reproduction steps, remediation, and a free retest, so the reports drop straight into your technical documentation and support your broader SOC 2, ISO 27001, and DORA programmes. Stingrai is a CREST-accredited penetration testing provider, founded in 2021, with offices in Toronto and London, which matters for buyers with EU exposure who want a credible tester on the record. See the full range on our services page, and engagement models on the pricing page.
Frequently asked questions
Does the EU AI Act require penetration testing of high-risk AI systems?
Not by name. Article 15 of the EU AI Act requires high-risk AI systems to be resilient against data poisoning, model poisoning, adversarial examples or model evasion, and confidentiality attacks or model flaws, and to resist unauthorised third parties exploiting system vulnerabilities. It never uses the words penetration testing. In practice, the way you demonstrate you addressed those attack classes is security testing evidence: adversarial robustness testing of the model plus penetration testing of the application and API around it, filed in your technical documentation.
What does Article 15 of the EU AI Act require?
Article 15 requires high-risk AI systems to achieve an appropriate level of accuracy, robustness, and cybersecurity and to perform consistently across those dimensions throughout their lifecycle. Paragraph 4 covers robustness, including technical redundancy, fail-safe plans, and mitigation of feedback loops in systems that keep learning. Paragraph 5 covers cybersecurity and names the specific AI attack classes a system must be resilient to. It sets outcomes rather than prescribing a particular test, which is why evidence of testing is how conformity is shown.
What attack classes does Article 15(5) name?
Article 15(5) names attacks that manipulate the training data set (data poisoning), attacks on pre-trained components used in training (model poisoning), inputs designed to cause the model to make a mistake (adversarial examples or model evasion), and confidentiality attacks or model flaws. Its opening sentence also requires resilience against unauthorised third parties who alter a system's use, outputs, or performance by exploiting vulnerabilities. Together these give you two testing surfaces: the model itself and the application around it.
When do the EU AI Act high-risk requirements apply?
After the Digital Omnibus, standalone high-risk systems listed in Annex III apply from 2 December 2027, and high-risk AI embedded as a safety component in products regulated under Annex I applies from 2 August 2028. These dates were deferred from the original 2 August 2026 and 2 August 2027. The Omnibus was adopted by the European Parliament on 16 June 2026 and approved by the Council on 29 June 2026; as of July 2026 it was signed and pending Official Journal publication, and the deferred dates take legal effect once published.
Did the Digital Omnibus change the Article 15 security requirements?
No. The Digital Omnibus deferred the application dates for high-risk AI obligations but did not change the substance of Article 15. The accuracy, robustness, and cybersecurity requirements, and the attack classes named in Article 15(5), are unchanged. The extra time is a grace window to build compliance, not a relaxation of what compliance means. Treat December 2027 as the date your evidence must be complete, not the date to start.
What is the difference between adversarial robustness testing and penetration testing?
Adversarial robustness testing targets the model's behaviour, using crafted inputs to try to flip an output, evade a safety filter, extract training data, or leak a system prompt, and it measures how often those attempts succeed. Penetration testing targets the application, API, and access controls around the model, finding flaws such as IDOR, broken authorisation, and business-logic abuse. Article 15(5) names both surfaces, so a complete evidence pack includes both, not one or the other.
What security testing evidence should I have ready before December 2027?
Aim for a report per named attack class: poisoning resilience, adversarial robustness, confidentiality attacks, model-flaw red teaming, and application and API penetration testing, each mapped to the Article 15 clause it evidences. Add severity and remediation records with retest results, coverage statements that show nothing material was skipped, and a retest plan that reflects Article 15's lifecycle duty. Start scoping now, because remediation and retesting take months, and clean evidence needs closed findings.
Who is responsible for Article 15 evidence, the provider or the deployer?
The provider, the entity that develops the high-risk system and places it on the market under its own name, carries the primary Article 15 duty and owns most of the testing and documentation. Deployers have lighter duties, but a deployer that renames the system, materially modifies it, or fine-tunes a model on its own data can inherit provider-level responsibilities and should procure its own testing for the parts it changed. If you buy a high-risk AI product, ask the provider for its Article 15 testing evidence.
References
European Union. Regulation (EU) 2024/1689 (EU AI Act), Article 15: Accuracy, robustness and cybersecurity. https://artificialintelligenceact.eu/article/15/. The operative text quoted throughout, including paragraph 5, which names data poisoning, model poisoning, adversarial examples or model evasion, and confidentiality attacks or model flaws.
Council of the EU. Artificial intelligence: Council gives final green light to simplify and streamline rules. 29 June 2026. https://www.consilium.europa.eu/en/press/press-releases/2026/06/29/artificial-intelligence-council-gives-final-green-light-to-simplify-and-streamline-rules/. Records the Council's final approval and the European Parliament's 16 June 2026 endorsement of the Digital Omnibus.
Gibson Dunn. EU AI Act Omnibus Agreement: Postponed High-Risk Deadlines and Other Key Changes. 2026. https://www.gibsondunn.com/eu-ai-act-omnibus-agreement-postponed-high-risk-deadlines-and-other-key-changes/. Confirms the deferred application dates: 2 December 2027 for standalone Annex III systems and 2 August 2028 for Annex I embedded product systems.
Freshfields. EU AI Act unpacked #34: The final Digital Omnibus on AI. 2026. https://www.freshfields.com/en/our-thinking/blogs/technology-quotient/eu-ai-act-unpacked-34-the-final-digital-omnibus-on-ai-key-amendments-to-the-a-102nber. Confirms the Omnibus deferred timelines only and left the underlying Article 15 obligations unchanged.
IBM. Cost of a Data Breach Report 2025. 30 July 2025. https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls. Source for the finding that 13% of organisations reported a breach of AI models or applications and 97% of those lacked proper AI access controls.



