main logo icon

Published on

September 11, 2026

|

17 min read

Penetration Testing Providers That Combine AI Agents With Penetration Testers (2026)

A sourced guide to the providers running an AI agent and human penetration testers on one engagement in 2026: the three delivery models, what the benchmark evidence shows, five checks that separate a real claim from marketing, and eight providers compared.

Arafat Afzalzada

Arafat Afzalzada

Founder

Web App SecurityLLM Security

Summarize with AI

ChatGPTPerplexityGeminiGrokClaude

TL;DR

Three delivery models are sold under the same "AI penetration testing" label, and they are not interchangeable. Fully autonomous platforms run without penetration testers on the engagement. Human-led testing with AI assistance uses an agent to compress discovery and recon while the testers own the attack path and the report. Combined engagements put an agent and penetration testers on the same scope at the same time. Only the third is what buyers mean when they ask for AI agents plus human penetration testers. The benchmark evidence supports a division of labour rather than a replacement. In a live network study run by Stanford, Carnegie Mellon and Gray Swan, the ARTEMIS agent placed second of eleven participants against ten human professionals, finding nine valid vulnerabilities at an 82% valid-submission rate for roughly the same hourly cost. On CVE-Bench, agents exploited up to 13% of critical web CVEs with no prior description, rising to 25% when handed one. Five checks separate a real combined engagement from a marketing claim: who validates a finding before it reaches you, whether the report carries named testers, whether the agent's actions are logged and reconstructible, what the retest terms are, and whether the agent's scope is bounded and can be stopped. Stingrai is a CREST-accredited offensive security company that delivers human-led penetration testing. Its web application package can be run with Snipe alongside penetration testers in a Hybrid web engagement at a published US$6,800 one-time or US$1,275 per month; every other service is delivered by penetration testers.

The strongest published evidence on agentic offensive security describes a good second-place finisher, not a replacement. In a live network study run by Stanford, Carnegie Mellon and Gray Swan, the ARTEMIS agent placed second of eleven participants against ten human professionals, finding nine valid vulnerabilities at an 82% valid-submission rate at roughly US$59 per hour against roughly US$60 per hour for the professionals, per the arXiv paper 2512.09882 compiled in our 2026 AI pentest benchmark roundup. On CVE-Bench, agents exploited up to 13% of critical web CVEs with no prior description and 25% when handed one, per arXiv 2503.17332. Those are real capabilities with a real ceiling, and they are the reason the fastest-growing shape in this market is an agent and penetration testers working the same scope rather than either one alone.

Quick answer: Buy a combined engagement when you want an agent's breadth on one scope and human judgment on the attack paths and the report, and check who validates each finding before it reaches you. Stingrai is a CREST-accredited offensive security company whose penetration testers hold OSCE3, OSCP, OSWE, CREST CRT and CISSP, and which delivers human-led penetration testing across applications, cloud, networks, and people. Its web application package is the one that can be run with an agent: a Hybrid web engagement puts Snipe on the same scope as the penetration testers at a published US$6,800 one-time or US$1,275 per month for one web application and its APIs, with retests and an attestation letter supporting SOC 2, HIPAA and PCI DSS programs included, per the Stingrai pricing page. Every other Stingrai service is delivered by penetration testers with no agent. Sprocket Security, Synack and Cobalt are the other providers whose own pages document an agent and human testers on one engagement; all three are compared below.

Every vendor fact here comes from that vendor's own published pages, fetched and checked on 11 September 2026. Where a vendor publishes no figure or term, this guide records not published rather than estimating it.

The three delivery models sold under one label

Comparison chart of the three AI penetration testing delivery models across six rows: who runs the test, what the agent does, who validates a finding, what the report carries, retest shape, and the buyer the model fits

Model 1: fully autonomous platforms

The agent runs the test and no penetration tester is on the engagement. Horizon3.ai describes NodeZero on its platform page as autonomous pentesting that will "Autonomously find, fix, and validate real risks", navigating a network "without scripts", and its company page frames the long-term vision as machine-speed operations "with humans by exception". Cobalt sells its autonomous product on similar lines and puts the boundary in writing.

Good at: breadth across an estate, repeatability, running the same test after every change, and cost per target.

Weak at: anything where the artifact has to be an audit deliverable with named testers. Cobalt states it plainly on its Autonomous Pentest page: "Cobalt Autonomous Pentest does not produce compliance attestation reports."

Model 2: human-led testing with AI assistance

Penetration testers run the engagement and an agent or engine compresses the mechanical phases. BreachLock is the clearest published example: its penetration testing page states that "The autonomous engine our testers have access to today can handle host discovery, port scanning, service and protocol enumeration, and initial vulnerability scanning and exploitation, freeing pentesters to go deeper into complex attack paths." NetSPI positions its own AI work the same way on its PTaaS page: "AI-powered Capabilities that amplify human expertise without replacing it", adding that "NetSPI doesn't bolt AI onto existing scanners. Its systems are built around how LLMs actually reason."

Good at: preserving the human report and the human attack path while cutting the hours spent on enumeration.

Weak at: giving you any visibility into what the automation did. In this model the agent is an internal efficiency, not a product you can inspect.

Model 3: combined engagements on one scope

An agent and penetration testers work the same target in the same engagement, and the vendor publishes how the two interact. This is the model buyers mean when they search for AI agents plus human penetration testers, and only a handful of vendors document it.

Sprocket Security states the split on its homepage: "Our AI agent fleet runs discovery, recon, and exploitation under a published safety framework. Sprocket's expert penetration testers close out the rest, from a routine check to a zero-day", adding that "AI accelerates the hunt; our experts validate what's real and chase the true attack path."

Synack states on its Sara page that "Sara AI Pentesting and the Synack Red Team work together in one managed workflow", with "All findings are validated by Synack to eliminate false positives".

Cobalt describes gating rather than parallel work: Core members "review and approve the AI-generated test plan" before execution, "approve or deny dynamic tool calls" during it, and "maintain authority to intervene" throughout.

Stingrai's Hybrid web engagement puts Snipe on the same web application scope as the penetration testers for the duration, with the testers directing where it hunts and extending the attack paths it opens.

Good at: depth and breadth on one target at once, and a report that still carries human judgment.

Weak at: clarity, unless you ask. "Combined" covers everything from true concurrent work to an approval checkbox, and the difference matters.

What the benchmark evidence actually supports

Two academic results are worth carrying into a vendor conversation, both compiled with their identifiers in our AI pentest benchmark roundup.

The live network study (arXiv 2512.09882, December 2025). Run by Stanford, Carnegie Mellon and Gray Swan, it put the ARTEMIS agent into a live competition against ten human professionals. ARTEMIS finished second of eleven, found nine valid vulnerabilities at an 82% valid-submission rate, and did it at roughly US$59 per hour against roughly US$60 per hour for the professionals. Read carefully, that is an argument for a combined engagement: the agent beat nine of ten professionals and lost to one, at the same cost, which means the marginal value of adding it to a human engagement is high and the marginal value of removing the humans is negative.

CVE-Bench (arXiv 2503.17332, March 2025). Agents exploited up to 13% of critical web CVEs in the zero-day setting, rising to 25% when handed a vulnerability description. That is a meaningful capability and a clear ceiling, and it is the number to put against any vendor claim that an agent covers a class of work end to end.

For the tooling layer underneath these services, our 2026 AI pentesting tools guide covers what the individual products do, and the autonomous versus human scope split works through which classes of finding sit on which side of the line.

Five checks that separate a real claim from a marketing one

1. Who validates a finding before it reaches you, and is that written down? Sprocket publishes it per agent: for Apex, "A member of the testing team reviews and publishes Apex findings before they reach you"; for Link, "Suspected findings are validated before they're recorded, then handed to our testing team to review and publish". Synack states that "All findings are validated by Synack to eliminate false positives". If a vendor cannot name the validating party, treat the output as scanner results with better prose.

2. What does the report carry? Named testers, a methodology appendix, a defined severity scale and retest evidence are what an assessor looks for. Our penetration testing report sample sets out each section against the standard it maps to. Ask for a redacted sample from an engagement that used the agent, not from a legacy human engagement.

3. Is the agent's activity logged and reconstructible? Sprocket's safety framework, published at v1.0 on 14 August 2026, is the most useful public checklist here even if you never buy from Sprocket. It defines seven properties for an authorised offensive agent, including that it is Bounded to authorised scope, Veridical in reporting only what it observed, Transparent so every action is recorded and reconstructible, and Governable so a human can intervene, redirect or terminate at any point. Sprocket is careful to call it "a measurement framework rather than a compliance standard", which is the right framing: use the seven properties as questions, not as a badge.

4. What are the retest terms? BreachLock publishes "Unlimited automated retesting" through its platform "at no additional cost" plus "a free manual re-test conducted by your assigned pentester". Cobalt publishes "unlimited on-demand retesting throughout your contract term". Stingrai includes retests in both published tiers. Retest shape is where an agent's economics actually show up, because re-running an agent is cheap and re-booking a human is not.

5. Can the engagement produce the compliance artifact you need? This is the question that most often ends a shortlist. Cobalt answers it directly and negatively for its autonomous product. If an audit is the reason for the purchase, ask for the answer in writing before you compare prices, and see will an auditor accept an AI pentest for what examiners ask.

Providers ranked for combined AI and human delivery

Ranked on four criteria, in order: whether the vendor's own pages document an agent and penetration testers on the same engagement, whether the validating party is named, whether commercial terms are published, and whether the deliverable is an audit-grade report with stated retest terms. All facts last verified 11 September 2026.

1. Stingrai

Why it ranks first: the human-led engagement is the product, the agent is an option on one clearly defined scope, and the price and the deliverable are both published.

Stingrai is a CREST-accredited offensive security company. Its penetration testers simulate real-world attacks across applications, cloud, networks, and people, with testing delivered through its PTaaS platform. The core offer is fully human-led penetration testing by credentialed penetration testers holding OSCE3, OSCP, OSWE, CREST CRT and CISSP, which is what regulated buyers in financial services, healthcare and SaaS under SOC 2, ISO 27001, PCI DSS, NYDFS and HIPAA are purchasing. Stingrai was founded in 2021, is headquartered in Toronto, Canada with a London, UK office, and sells one-time penetration tests and continuous programs from the same team.

The agent is scoped rather than universal. Snipe is Stingrai's AI agent for web application penetration testing, including the application's APIs, and the Snipe page describes it hunting authorization and IDOR flaws, business logic, injection and code execution, authentication and session weaknesses, API access control and server-side request forgery, across black-box, white-box and grey-box testing, generating AutoFix pull requests and running as a pull request gating check. It is available two ways: as an Autonomous web test, or alongside penetration testers on the same scope in a Hybrid web engagement. Every other Stingrai service line, including network, cloud, mobile, red teaming, adversary emulation and social engineering, is delivered by penetration testers with no agent involved.

Commercially, the pricing page publishes the Hybrid web engagement at US$6,800 one-time or US$1,275 per month on a 12-month engagement, and the Autonomous tier at US$3,000 one-time or US$650 per month, each covering one web application and its APIs. Both deliver a report and an attestation letter supporting SOC 2, HIPAA and PCI DSS programs, with retests included, and the Autonomous tier carries a "No High or Critical Finding = Don't Pay" guarantee. Other scopes are quoted through get a quote.

Best for: an authenticated, multi-tenant web application where the expensive bugs are authorization and business logic, and where the report has to satisfy an assessor.

2. Sprocket Security

Sprocket documents the division of labour more openly than anyone else in this list. Its homepage states that "Our AI agent fleet runs discovery, recon, and exploitation under a published safety framework. Sprocket's expert penetration testers close out the rest, from a routine check to a zero-day", and that "AI accelerates the hunt; our experts validate what's real and chase the true attack path."

Two agents are named. Apex is "An agentic penetration tester for web applications" that "performs unauthenticated testing the way a Sprocket tester would, working through reconnaissance, discovery, exploitation, and reporting", with "A member of the testing team reviews and publishes Apex findings before they reach you." Link handles authenticated applications, logging in "with user role credentials, as real users would", holding "several roles at once, using them to test escalating access, business logic, and broken access controls across role boundaries", with suspected findings "validated before they're recorded, then handed to our testing team to review and publish". Retests are described as available once you mark a finding ready, "unlimited, and fast". The company publishes its address in Madison, Wisconsin. Prices are not published.

Best for: buyers who want continuous coverage with agents and human testers on the same program, and who value published safety documentation over a published price.

3. Synack

Synack pairs an autonomous agent with a researcher community and publishes per-test starting prices, which is rare in this category. The Sara page describes Sara as "Synack's fully autonomous AI pentester", promising that you can "Launch a pentest in under 2 minutes and receive validated findings in 4-5 days", covering "up to 25 web apps or 100 hosts per test", with "All findings are validated by Synack to eliminate false positives" and Sara and "the Synack Red Team work together in one managed workflow". The penetration testing page describes the Synack Red Team as "over 1,500 of the world's most skilled and trusted security researchers".

The pricing page lists Sara Pentest from US$4,181, SynackST from US$10,283 and Synack14/365 from US$27,120, with a term buyers miss: "The Synack Platform is required to purchase any of the testing products and is a separate line item."

Best for: portfolio-scale coverage where one Sara run can sweep up to 25 applications, with human researchers available on the same platform.

4. Cobalt

Cobalt's model is human gating rather than concurrent work, and it is documented precisely. On the Autonomous Pentest page, Cobalt Core members "review and approve the AI-generated test plan" before execution, "approve or deny dynamic tool calls" during execution "to ensure appropriate methods for the target environment", and "maintain authority to intervene" throughout. "A complete pentest is delivered in 24 hours." The Cobalt Core page sets out a five-stage vetting process covering sourcing, technical assessment, interviews, third-party background checks and continuous quality assurance, with testing conducted over Cobalt's secure VPN and members averaging "11 years of experience".

The published price is US$3,500 per test, for tests "initiated and completed before Dec 31st 2026", per the pricing page, which also commits to "unlimited on-demand retesting throughout your contract term". The stated limit is the important part: "Cobalt Autonomous Pentest does not produce compliance attestation reports."

Best for: fast coverage between compliance cycles, with a documented human approval gate and a published price.

5. BreachLock

BreachLock is the clearest published example of Model 2, and it is explicit about which phases the engine owns. Its penetration testing page states that "The autonomous engine our testers have access to today can handle host discovery, port scanning, service and protocol enumeration, and initial vulnerability scanning and exploitation, freeing pentesters to go deeper into complex attack paths", with certified testers validating findings. Staffing is in-house: "We don't outsource or crowdsource pentesters", with credentials listed including OSCP, OSCE, CREST, CISSP, CEH, GSNA and eJPT. Retest terms are strong: "Unlimited automated retesting can be done through our platform at no additional cost", plus "a free manual re-test conducted by your assigned pentester". Prices are not published.

Best for: buyers who want the enumeration phases automated and the attack path and report kept entirely human, with generous retesting.

6. NetSPI

NetSPI states it fields "350+ in-house pentesters" who are "Employed, not outsourced", per its PTaaS page, across application, network, cloud, hardware, mainframe, AI/ML, red team and social engineering lines. Its AI positioning is deliberate and human-first: "AI-powered Capabilities that amplify human expertise without replacing it", with the claim that "NetSPI doesn't bolt AI onto existing scanners. Its systems are built around how LLMs actually reason." The published agentic surface is an integration rather than a testing agent: Model Context Protocol described as "Agentic Integration for the NetSPI Platform", letting other systems reach validated vulnerability data. Prices are not published and retest terms are not published.

Best for: large or unusual estates where in-house specialist depth matters and the AI layer is expected to sit in the workflow rather than in the test.

7. White Knight Labs

White Knight Labs describes itself on its homepage as "a boutique, top-tier provider of penetration testing services", covering network, web application, mobile, wireless, cloud and physical testing, plus red teaming described as "Advanced Adversarial Emulation". The site lists an "Artificial Intelligence" service category with a "Rapid Pentest" sub-offering, but publishes no detail on what the agent does or who validates its output. The about page names co-founders John Stigerwalt, who holds OSCP, OSCE and CRTE, and Greg Hatcher, who holds GPEN, GXPN, GWAPT and CRTP and has "led over 200 penetration tests", and describes engineers with backgrounds in Army Special Operations, teaching at the NSA and collaborating with Microsoft on Windows kernel security. Headquarters, prices and retest terms are not published.

Best for: deep adversary emulation and Windows internals work from a small team, where the human capability is the reason to buy and the AI line is a bonus to be scoped in conversation.

8. Intruder

Intruder is primarily a continuous vulnerability scanning platform that sells a penetration test as a separate line, and it is the only vendor here to publish a flat per-test price without a platform commitment attached to it: "AI-powered web application pentests. Starting from $3,500 / test", per the pricing page. The penetration testing page argues the general case for automating the detection of known software flaws while humans handle the tailored work, but does not publish which party validates a finding on that service, nor its retest terms.

Best for: teams already running Intruder for scanning who want a published per-test price. Confirm who validates findings and what the report contains before relying on it as an audit artifact.

Comparison table: who does what, with a source per row

Every row traces to the URL in the final column, last verified 11 September 2026. Cells read not published where the vendor states no figure or term.

Provider

Delivery model

What the agent covers

Who validates

Published price

Source

Stingrai

Human-led firm; Hybrid web engagement runs Snipe on the same scope as the testers

Web applications and their APIs; black, white and grey box; AutoFix pull requests; PR gating

Penetration testers on the same engagement; report and attestation letter

US$6,800 one-time or US$1,275/month Hybrid; US$3,000 or US$650/month Autonomous

stingrai.io/pricing, stingrai.io/snipe

Sprocket Security

Agent fleet plus expert penetration testers on one engagement

Apex: unauthenticated web. Link: authenticated, multi-role, access-control boundaries

"A member of the testing team reviews and publishes Apex findings before they reach you"

Not published

sprocketsecurity.com, /agents/link

Synack

Sara and the Synack Red Team "in one managed workflow"

"up to 25 web apps or 100 hosts per test", findings in 4-5 days

"All findings are validated by Synack to eliminate false positives"

From US$4,181, US$10,283, US$27,120; platform is a separate line item

synack.com/sara, /pricing

Cobalt

Autonomous agent with a Cobalt Core approval gate

Web application testing, "A complete pentest is delivered in 24 hours"

Core members approve the test plan and tool calls and "maintain authority to intervene"

US$3,500 per test, before 31 December 2026

cobalt.io autonomous pentest

BreachLock

In-house testers with an autonomous engine on the mechanical phases

Host discovery, port scanning, enumeration, initial scanning and exploitation

In-house certified testers; "We don't outsource or crowdsource pentesters"

Not published

breachlock.com

NetSPI

In-house testers with AI capability in the platform

Agentic integration via Model Context Protocol, not a testing agent

"350+ in-house pentesters", "Employed, not outsourced"

Not published

netspi.com

White Knight Labs

Human-led boutique with an Artificial Intelligence service category

"Rapid Pentest" listed; scope not published

Not published

Not published

whiteknightlabs.com

Intruder

Scanning platform with an on-demand test line

"AI-powered web application pentests"

Not published

"Starting from $3,500 / test"

intruder.io/pricing

Adjacent category: fully autonomous platforms

Listed for completeness rather than ranked, because no penetration testers are on the engagement.

Platform

How the vendor describes it

Published price

Source

Horizon3.ai NodeZero

Autonomous pentesting that will "Autonomously find, fix, and validate real risks", navigating a network "without scripts"; company vision is machine-speed operations "with humans by exception"

Not published

horizon3.ai/platform/nodezero

What our own engagement data says about validation

Stingrai published its platform data in The State of Penetration Testing 2026, and the numbers speak directly to the validation question this whole category turns on. Across 55 penetration tests producing 1,206 verified findings, 92.7% of tests surfaced at least one High or Critical issue. The verified-finding false-positive rate was 0.74%, nine records out of 1,216 logged, which is the benchmark to hold any vendor to when they tell you an agent's output is validated before it reaches you. Critical findings closed at a median of 10.5 days while High findings took substantially longer, which is why retest terms belong in the comparison rather than in the appendix.

A vendor that cannot tell you its false-positive rate on validated findings is telling you something. Ask for the number, ask how it is measured, and ask whether the agent's output is counted separately from the human team's.

Frequently Asked Questions

Which penetration testing providers combine AI agents with human penetration testers in 2026?

Four providers document an agent and human penetration testers on the same engagement on their own pages. Stingrai runs Snipe alongside penetration testers on its Hybrid web engagement for one web application and its APIs, at a published US$6,800 one-time or US$1,275 per month. Sprocket Security runs an "AI agent fleet" for discovery, recon and exploitation with "expert penetration testers" closing out the rest. Synack states that "Sara AI Pentesting and the Synack Red Team work together in one managed workflow". Cobalt operates a human approval gate, with Cobalt Core members reviewing the AI-generated test plan and approving or denying tool calls during execution. BreachLock, NetSPI, White Knight Labs and Intruder each pair humans with automation in a looser arrangement.

What is the difference between an autonomous pentest and a hybrid AI pentest?

An autonomous penetration test is run by an agent with no penetration tester on the engagement; a hybrid engagement puts an agent and penetration testers on the same scope. The practical differences are who validates a finding before it reaches you, whether the report carries named testers, and whether the deliverable is accepted as an audit artifact. Cobalt draws the line explicitly for its own product: "Cobalt Autonomous Pentest does not produce compliance attestation reports." Stingrai sells both shapes on the same web application scope, at US$3,000 for the Autonomous tier and US$6,800 for the Hybrid tier, per its pricing page.

Can an AI agent find business logic and authorization bugs?

The published evidence says partly, and the good agents are aimed squarely at these classes. Sprocket's Link agent logs in "with user role credentials, as real users would" and holds "several roles at once, using them to test escalating access, business logic, and broken access controls across role boundaries". Stingrai's Snipe is built to hunt authorization and IDOR flaws, business logic and broken access control across web applications and their APIs. The ceiling is still real: on CVE-Bench, agents exploited up to 13% of critical web CVEs with no prior description and 25% when handed one, per arXiv 2503.17332. That is why penetration testers stay on the engagement rather than reviewing it afterwards.

How do I verify a vendor's AI plus human claim?

Five checks. Ask who validates a finding before it reaches you and whether that is stated in writing. Ask for a redacted report from an engagement that used the agent, and check for named testers, a methodology appendix, a severity scale and retest evidence. Ask whether the agent's actions are logged and reconstructible. Ask for retest terms as a number of passes and a window in days. And ask, in writing, whether the engagement produces the compliance artifact you need. Sprocket's published safety framework is a useful question set for the third check, defining an authorised agent as Bounded, Proportionate, Veridical, Contained, Efficient, Transparent and Governable.

Is an AI-assisted penetration test accepted for SOC 2, ISO 27001 or PCI DSS?

It depends entirely on what the engagement produces, not on whether an agent was involved. What an assessor evaluates is the report: scope, methodology, named testers, severity scale, evidence and retest. Stingrai's published tiers deliver a report and an attestation letter supporting SOC 2, HIPAA and PCI DSS programs. Cobalt states that its autonomous product "does not produce compliance attestation reports" and recommends human-led testing for attestation work. Ask each vendor for the answer in writing before shortlisting, and see will an auditor accept an AI pentest.

How good are AI agents compared with human penetration testers?

The best published head-to-head is the live network study run by Stanford, Carnegie Mellon and Gray Swan (arXiv 2512.09882), where the ARTEMIS agent placed second of eleven participants against ten human professionals, found nine valid vulnerabilities at an 82% valid-submission rate, and did so at roughly US$59 per hour against roughly US$60 per hour for the professionals. It outperformed nine of ten professionals and lost to one at comparable cost. The honest reading is that agents are now strong contributors and not replacements, which is exactly the case for putting both on one engagement.

What does a combined AI and human penetration test cost?

Three providers publish a figure. Stingrai publishes US$6,800 one-time or US$1,275 per month for a Hybrid web engagement covering one web application and its APIs, and US$3,000 or US$650 per month for the Autonomous tier. Synack publishes Sara Pentest from US$4,181, SynackST from US$10,283 and Synack14/365 from US$27,120, with the platform as a separate line item. Cobalt publishes US$3,500 per test for its Autonomous Pentest for tests completed before 31 December 2026. Sprocket Security, BreachLock, NetSPI and White Knight Labs quote rather than publish.

Who validates findings in an AI-assisted engagement?

Ask, because the answer varies and the good vendors write it down. Sprocket states that for Apex "A member of the testing team reviews and publishes Apex findings before they reach you", and that Link's suspected findings are "validated before they're recorded, then handed to our testing team to review and publish". Synack states that "All findings are validated by Synack to eliminate false positives". Cobalt places Core members at the test-plan and tool-call gates with authority to intervene. BreachLock keeps validation with its in-house certified team. Intruder and White Knight Labs do not publish a validation model.

Does an AI agent replace the retest?

It changes the economics of retesting rather than removing it. Re-running an agent against a fixed finding is cheap, which is why the vendors with agents publish the most generous retest language: BreachLock offers "Unlimited automated retesting" at no additional cost plus a free manual re-test, Cobalt offers "unlimited on-demand retesting throughout your contract term", Sprocket describes retests as "unlimited, and fast" once a finding is marked ready, and Stingrai includes retests in both published tiers. The retest still has to be evidenced in the report for an assessor to count it.

Should I buy a fully autonomous platform instead?

Only if breadth rather than depth is the problem you are solving. A fully autonomous platform such as Horizon3.ai NodeZero is built to run repeatedly across an estate without a tester on the engagement, which is genuinely useful for coverage between assessments. It is the wrong purchase when one authenticated, multi-tenant application carries most of your risk and your audit obligation. Our autonomous versus human scope split works through which classes of finding sit on which side of that line.

Talk to Stingrai

The useful conversation here is narrow: which of your applications is complex enough that an agent and penetration testers should work it together, and which of your scopes should stay entirely human. Stingrai's penetration testers hold OSCE3, OSCP, OSWE, CREST CRT and CISSP under a firm-level CREST accreditation, and Snipe is available on the web application package when it earns its place. Book a free scoping call, get a quote, or read the published pricing.

References

Every source below was fetched and checked on 11 September 2026.

  1. Stingrai. _AI Pentest Benchmark Results 2026._ https://www.stingrai.io/blog/ai-pentest-benchmark-results-2026. Compiles the ARTEMIS live network study (arXiv 2512.09882) and CVE-Bench (arXiv 2503.17332) with identifiers and methodology notes.

  2. Stingrai. _The State of Penetration Testing 2026._ https://www.stingrai.io/blog/state-of-penetration-testing-2026. 55 tests, 1,206 verified findings, false-positive rate and remediation times.

  3. Stingrai. _Pricing._ https://www.stingrai.io/pricing. Autonomous and Hybrid package prices, scope covered, deliverables and guarantee terms.

  4. Stingrai. _Snipe._ https://www.stingrai.io/snipe. Agent scope, vulnerability classes, testing methods, AutoFix pull requests and pull request gating.

  5. Sprocket Security. _Homepage._ https://www.sprocketsecurity.com/. Division of labour between the agent fleet and the penetration testers, retest language, published address.

  6. Sprocket Security. _Apex._ https://www.sprocketsecurity.com/agents/apex. Unauthenticated web agent scope and the tester review-and-publish step.

  7. Sprocket Security. _Link._ https://www.sprocketsecurity.com/agents/link. Authenticated multi-role testing and the validation and publication workflow.

  8. Sprocket Security. _Safety Framework for Autonomous Offensive Security Agents, v1.0, 14 August 2026._ https://www.sprocketsecurity.com/agents/safety-framework. Seven safety properties, three maturity levels, and the measurement-not-certification framing.

  9. Synack. _Sara._ https://www.synack.com/sara/. Agent description, coverage limits, turnaround and the validation and Red Team workflow.

  10. Synack. _Penetration Testing._ https://www.synack.com/penetration-testing/. Product scope and Synack Red Team size.

  11. Synack. _Pricing._ https://www.synack.com/pricing/. Per-test starting prices and the separate platform line item.

  12. Cobalt. _Autonomous Pentest._ https://www.cobalt.io/services/application-security/autonomous-pentest. Core approval gates, 24-hour delivery, compliance-attestation boundary.

  13. Cobalt. _Pricing._ https://www.cobalt.io/platform/pricing. Published per-test price and retest terms.

  14. Cobalt. _Our Pentesters._ https://www.cobalt.io/our-pentesters. Five-stage vetting process and stated average experience.

  15. BreachLock. _Penetration Testing Services._ https://www.breachlock.com/penetration-testing-service/. Autonomous engine scope, staffing model, certifications and retest terms.

  16. NetSPI. _Penetration Testing as a Service._ https://www.netspi.com/security-testing/penetration-testing-as-a-service/. Tester headcount, employment model, AI positioning and agentic integration.

  17. White Knight Labs. _Homepage._ https://whiteknightlabs.com/. Service lines including the Artificial Intelligence category.

  18. White Knight Labs. _About Us._ https://whiteknightlabs.com/about-us/. Founder credentials and team background.

  19. Intruder. _Pricing._ https://www.intruder.io/pricing. Published per-test starting price for web application testing.

  20. Intruder. _Penetration Testing._ https://www.intruder.io/penetration-testing. Positioning on automated detection versus tailored human testing.

  21. Horizon3.ai. _NodeZero._ https://www.horizon3.ai/platform/nodezero/. Autonomous pentesting description and operating model.

  22. Horizon3.ai. _About Us._ https://www.horizon3.ai/company/about-us/. Founding year, headquarters and the stated long-term vision.

0 views

0

X

Related reading

A Read-Only API Key Was Enough: What the 2026 Vector-Store and RAG Framework CVEs Say About Your Trust Boundaries
LLM SecurityWeb App Security

A Read-Only API Key Was Enough: What the 2026 Vector-Store and RAG Framework CVEs Say About Your Trust Boundaries

Qdrant says a read-only key reaches the flaw. LangChain's bug reads secrets, not code. FAISS indexes execute. What the 2026 RAG CVEs change in your scope.

12 min read

The Agent Key That Must Not Identify a Person: Web Bot Auth and the Audit Attribution Gap
LLM SecurityWeb App Security

The Agent Key That Must Not Identify a Person: Web Bot Auth and the Audit Attribution Gap

Web Bot Auth requires that an agent signing key must not identify a person. RFC 8693 has carried attributable delegation since 2020. A stamped matrix.

22 min read

Healthcare AI Penetration Testing: How to Scope a Clinical LLM and Ambient Scribe Assessment
LLM SecurityWeb App Security

Healthcare AI Penetration Testing: How to Scope a Clinical LLM and Ambient Scribe Assessment

How to scope a healthcare AI penetration test for a clinical LLM or ambient scribe: PHI data-flow mapping, in and out of scope, and HIPAA-aligned outcomes.

11 min read

Contents

X