main logo icon

Published on

July 22, 2026

|

16 min read

The Pentest and Red Team RFP Question Bank: 75 Scored Questions With Red Flag Answers

A scored penetration testing and red team RFP question bank: 75 vendor questions across 7 weighted sections, each with a strong-answer and red-flag key, plus AI-disclosure and DORA TLPT tester-qualification sections.

Arafat Afzalzada

Arafat Afzalzada

Founder

Web App SecurityNetwork Security

Summarize with AI

ChatGPTPerplexityGeminiGrokClaude

TL;DR

A penetration testing or red team RFP should score vendor answers, not just collect them. This bank gives you 75 questions across seven weighted sections that total 100: general and methodology, tester qualifications and red team, cloud and infrastructure depth, AI and automation disclosure, reporting quality, retesting and remediation, and data handling. Each of the strongest 30 questions ships with a red-flag answer key so your evaluators know what a 5 sounds like and what a 0 sounds like before the calls start. Two sections most templates skip carry the most weight this year. AI disclosure matters because support for fully automated pentesting fell from 29 percent to 9 percent in a single year and 78 percent of teams hit critical false negatives from automated tools (Cobalt, 2026). Tester qualifications matter because regulated red team work already has hard bars: the DORA threat-led testing rules require a lead with at least five years of experience, and the TIBER-EU and CBEST schemes set similar thresholds. Score every bidder on the same axes, treat a scanner sold as a pentest as an automatic red flag, and let the total decide.

A penetration testing RFP is only as strong as the answers it lets you compare. Most templates hand you a list of questions and no way to grade the replies, so every vendor reads as competent and the cheapest bid wins by default. This bank is built the other way around. It gives you 75 vendor questions across seven weighted sections that total 100, a scored answer key that spells out what a strong reply contains and what a red-flag reply sounds like, a dedicated AI and automation disclosure section, and tester-qualification questions drawn straight from the regulated red team frameworks. Those four things, weighted scoring, a red-flag key, AI disclosure, and regulator-grade tester questions, are exactly what a generic template leaves out.

The core question this answers: what questions should you include in a penetration testing, red team, or cloud pentest RFP, and which vendor answers are red flags? Ask about methodology, tester qualifications, cloud depth, AI and automation disclosure, reporting quality, retesting, and data handling. Treat these replies as red flags: a scanner-only workflow sold as a pentest, testers the vendor will not name or credential, no sample report, findings with no retest, and any refusal to say which findings a human actually validated. Every one of those is scored below, so the flag turns into a number rather than a gut feeling.

This is written for the procurement lead, security engineer, or CISO who is drafting a vendor questionnaire right now and needs the replies to be comparable when they come back. Copy the questions into your RFP, set the weights to your own risk, and hand your evaluators the answer key so three people scoring the same proposal land within a point of each other. If you are earlier in the process and still comparing prices, start with how to compare penetration testing quotes, then come back here to score the technical response.

Key takeaways

  • Score the answers, do not just collect them. A template that gathers replies without a rubric lets every vendor sound equal. Weighting seven sections to 100 and scoring each answer 0 to 5 turns a stack of proposals into a ranking your whole committee can defend.

  • AI disclosure is the new must-ask section. Support for fully automated pentesting fell from 29 percent to 9 percent in a single year, and 78 percent of teams reported fully automated tools missing critical vulnerabilities and returning false negatives (Cobalt, 2026). An RFP that does not ask what runs the test is buying blind.

  • Regulated red team work already sets a tester-qualification bar. The DORA threat-led penetration testing rules require a lead with at least five years of penetration and red team experience plus two more testers with at least two years each (EU 2025/1190, Article 7). You can lift that bar into any red team RFP as pass-or-fail criteria.

  • A scanner sold as a pentest is the most common red flag. The cheapest way to lose the value of a test is to buy an automated scan dressed up as manual testing. Every methodology and reporting question below is designed to surface that substitution before you sign.

  • Retesting and data handling are where thin bids hide cost. A proposal that excludes fix verification, or that is vague about where your data lives and how long it is kept, is cheaper today and more expensive the first time you need evidence for an audit or an incident.

Methodology and sources

This bank is a practitioner tool, not a survey, but the numbers and thresholds inside it are sourced to named primary publishers so you can cite them in your own RFP:

  • Cobalt, AI and Pentesting Pulse Report 2026 (June 2026): a survey of 455 security leaders and practitioners built on five years of pentesting data. Source for the year-over-year drop in support for fully automated pentesting (29 percent to 9 percent), the 78 percent critical-false-negative figure, and the 47 percent hybrid-model preference.

  • HackerOne, 2025 Hacker-Powered Security Report (November 2025): source for the 210 percent year-over-year jump in valid AI-assisted vulnerability reports.

  • DORA TLPT Regulatory Technical Standards, Commission Delegated Regulation (EU) 2025/1190 (applicable 8 July 2025): source for the threat-led penetration testing tester-qualification requirements in Article 7.

  • TIBER-EU framework and service provider procurement guidance, European Central Bank: source for the red team manager and red team member experience thresholds.

  • CBEST implementation guide, Bank of England, and CREST: source for the CREST accreditation and certification expectations that govern intelligence-led financial-sector testing.

The research cutoff for this pass was July 2026. Where a figure could not be confirmed against its named primary publisher on this pass, it was left out rather than estimated. Every stat in the sections below links back to its publisher so any claim can be audited inline.

How to use this bank

Rfp Bank Scoring Model

The seven sections carry a default weight that sums to 100. Adjust the weights to your own risk, a cloud-native company might push cloud depth to 20 and trim red team, a regulated bank will do the reverse, but set the weights before you read a single proposal and keep them identical across every bidder. Changing weights mid-evaluation is how a committee talks itself into the vendor it already liked.

Score each answer on a 0 to 5 scale: 0 for a red-flag answer, 3 for adequate, 5 for the strong answer described in the key. Multiply each section's average by its weight, then total to 100. Add one override rule that protects the whole exercise: a red-flag answer on any high-weight question caps that section, no matter how polished the rest of the reply reads. A vendor that cannot name its testers does not get to make it up on price.

The strongest 30 questions below ship with a full answer key in table form. The remaining questions are listed compactly by section so you can drop the whole bank into your RFP and still keep the document readable. If you want a scope document to sit alongside the questionnaire, the AI and LLM pentest scope-of-work template pairs cleanly with this bank.

The red-flag answer key, explained

Rfp Bank Red Flag Anatomy

Each scored question has three parts: the question you ask, the signals a strong answer contains, and the red-flag answer that caps or fails the score. The point of writing the red flag down in advance is consistency. When three evaluators independently know that "we run an automated scanner and pass you the output" is a 0 on a manual-methodology question, they score the same proposal the same way. Without the key, one evaluator hears "we use automation for efficiency" as innovative and another hears it as a substitution, and your ranking becomes an argument about tone instead of substance.

Section 1: General and methodology (weight 15)

This section separates a structured, standards-based engagement from an ad hoc scan. You are looking for a named methodology, manual testing of the classes automated tools miss, and a scope conversation that starts with your risk rather than a fixed price list.

Question

A strong answer contains

Red-flag answer

1. What methodology and standards do you follow, and how do you adapt them per engagement?

Named framework (OWASP WSTG, PTES, NIST SP 800-115, OSSTMM) applied as a baseline, then tailored to the target's architecture and threat model with examples.

"We use our own proprietary process" with no named standard and no detail, or a pure tool list.

2. What share of the test is manual versus automated, and which findings come from which?

A clear split, with manual effort focused on authorization, business logic, and chained exploits, and automation limited to coverage and enumeration.

"It is fully automated" or an inability to say which findings a human produced.

3. How do you scope authorization, IDOR, and business-logic flaws that a scanner cannot see?

Concrete manual test cases per role and object, multi-account testing, and worked examples from a sample report.

Business logic treated as out of scope, or folded into a scanner run with no manual cases.

4. How do you define and justify severity, and do you use a named rating system?

CVSS or a documented internal model tied to real business impact and exploitability, not raw scanner severities.

Severities copied straight from a tool with no impact analysis.

5. Who owns the engagement day to day, and how do we reach the testers, not just an account manager?

A named technical lead, a defined communication channel, and direct tester access during the test window.

All contact routed through sales, with no access to the people doing the work.

Remaining general and methodology questions to include, scored on the same 0 to 5 scale:

  1. What certifications and accreditations does your firm hold, and which are firm-level versus individual?

  2. How long has the firm been delivering this exact type of testing, and can you share relevant references?

  3. How do you handle scope changes discovered mid-test, and what is the change-control process?

  4. What is your typical timeline from kickoff to draft report for an engagement of our size?

  5. How do you avoid and disclose conflicts of interest, including reselling the products you assess?

  6. What does your kickoff require from us, and how do you keep the engagement from stalling on our side?

  7. How do you measure the quality of your own testing, and what stops a weak test shipping as a strong report?

Section 2: Tester qualifications and red team (weight 20)

Rfp Bank Tester Quals

This is the highest-weighted section because the people matter more than the process. A named, credentialed tester with years of relevant experience finds the bugs that decide the engagement. This section also carries the regulator-grade questions: if you are a financial entity in scope for the EU's Digital Operational Resilience Act, or you run intelligence-led testing under the UK's CBEST scheme, the tester bar is not a preference, it is written into the rules.

The DORA threat-led penetration testing Regulatory Technical Standards, applicable since 8 July 2025, require external testers to field a lead with at least five years of penetration and red team testing experience, plus at least two more testers with at least two years each, a combined five or more prior assignments, at least five references, curricula vitae with market-standard certifications, and professional indemnity insurance (EU 2025/1190, Article 7). The ECB's TIBER-EU procurement guidance sets a parallel bar: a red team manager with at least five years of experience including three years leading tests in financial services, and each red team member with at least two years (TIBER-EU, ECB). Any buyer can lift these thresholds into an RFP as pass-or-fail criteria, even outside financial services.

Question

A strong answer contains

Red-flag answer

1. Who specifically will test our environment, and what are their qualifications?

Named testers with relevant certifications (OSCP, OSCE3, OSWE, CREST CRT, CRTO) and years of hands-on experience on similar targets.

"We will assign qualified staff" with no names, no certifications, and no CVs.

2. Are the testers your own staff or subcontractors, and where are they located?

In-house testers, or disclosed and vetted subcontractors, with locations relevant to your data-residency needs.

Undisclosed subcontracting, or refusal to say who actually performs the work.

3. For a red team engagement, what is your adversary-emulation methodology and how do you map to MITRE ATT&CK?

Threat-intelligence-led scenarios, explicit ATT&CK technique mapping, and objectives tied to your crown-jewel assets.

A vulnerability scan relabeled as a red team, with no adversary model or ATT&CK mapping.

4. For regulated testing (DORA TLPT, TIBER-EU, CBEST), how do your testers meet the scheme's qualification requirements?

Specific mapping to the scheme: lead experience years, references, certifications, insurance, and separation of threat intelligence from red team roles.

Vague "we are compliant" with no reference to the actual tester-qualification thresholds.

5. How do you separate the threat-intelligence function from the red team, and why does that matter?

Independent intelligence provider or a walled-off internal team, so scenarios reflect real adversaries rather than the red team's convenience.

Threat intelligence and red team collapsed into one role with no independence.

6. What are the rules of engagement, deconfliction, and stop conditions for a live red team test?

A signed rules-of-engagement document, named deconfliction contacts, immediate stop authority, and legal authorization to test.

No written rules of engagement, or no clear way to halt the operation.

Remaining tester and red team questions to include:

  1. What is the average tenure and seniority of the testers you will assign to us?

  2. How do you keep testers current: continuing certification, internal research time, CVE publication, conference work?

  3. Can we interview or veto the proposed lead tester before the engagement starts?

  4. For physical or social-engineering components, what authorization and safety controls do you require in writing?

  5. How do you handle a tester leaving mid-engagement without losing continuity or context?

  6. What is your process for responsible disclosure if you find a zero-day in third-party software during our test?

  7. How do you demonstrate independence from the vendors and products in our environment?

  8. What insurance and liability cover do you carry, and can you provide a certificate?

For grading the red team portion in depth, pair this section with grading MITRE ATT&CK coverage in a red team proposal and the red team rules-of-engagement buyer checklist.

Section 3: Cloud and infrastructure depth (weight 15)

Cloud RFP criteria are where generic pentest shops thin out fastest. Testing an AWS, Azure, or GCP environment properly means reviewing identity and access management, control-plane configuration, and workload isolation, not just running a network scan against public IPs. This section is worth 15 because a shallow cloud test gives false comfort in exactly the layer where most modern breaches now start.

Question

A strong answer contains

Red-flag answer

1. How do you test cloud IAM: roles, trust policies, privilege escalation paths, and cross-account access?

Manual IAM policy review plus escalation-path testing (for example, role chaining and confused-deputy issues), with named cloud experience.

"We scan the external IP range" with no IAM or control-plane testing at all.

2. How do you assess the cloud control plane and configuration, not just the workloads?

Review of the management plane against a named benchmark (CIS, provider well-architected guidance), with misconfiguration testing.

Configuration review skipped, or reduced to a generic checklist with no manual validation.

3. How do you test container and Kubernetes security: escapes, RBAC, secrets, and workload isolation?

Concrete cluster test cases: pod escape attempts, RBAC review, secrets handling, and namespace isolation checks.

Kubernetes treated as out of scope or "the same as any Linux host."

4. What access model do you use for cloud testing, and how do you protect the credentials we provide?

Least-privilege scoped roles, short-lived credentials, and a clear handling and destruction process for anything we issue.

Requests for long-lived admin keys with no explanation of how they are protected or revoked.

Remaining cloud and infrastructure questions to include:

  1. How do you test serverless functions, managed services, and provider-specific attack surface?

  2. How do you approach hybrid and on-premise-to-cloud trust boundaries?

  3. Do you test the CI/CD pipeline and infrastructure-as-code, and how?

  4. How do you avoid triggering cloud-provider abuse detection or disrupting shared tenancy?

  5. What cloud-specific certifications or hands-on experience do your assigned testers hold?

  6. How do you validate network segmentation and lateral movement inside a virtual private cloud?

  7. How do you handle testing across multiple accounts, subscriptions, or projects in one engagement?

Section 4: AI and automation disclosure (weight 15)

Rfp Bank Ai Disclosure

This section did not exist in most RFP templates a year ago, and it is now non-negotiable. In Cobalt's AI and Pentesting Pulse Report 2026, a survey of 455 security leaders and practitioners, support for fully automated pentesting fell from 29 percent to 9 percent year over year, 78 percent of organizations reported fully automated scanning tools missing critical vulnerabilities and returning false negatives, and 47 percent now prefer a hybrid model where humans support AI testing (Cobalt, 2026). Adoption of AI in testing is still climbing, HackerOne recorded a 210 percent year-over-year jump in valid AI-assisted vulnerability reports (HackerOne, 2025), so the question is not whether a vendor uses AI, it is whether they disclose it and validate it.

Question

A strong answer contains

Red-flag answer

1. Do you use AI or automated agents in testing, and which findings come from them versus a human?

Transparent disclosure of where AI is used, with human validation of every high-severity finding before it reaches your report.

"No" that later turns out to be untrue, or "everything is human" with no way to verify.

2. How do you prevent AI-driven false positives from reaching our report?

A named validation step: a senior tester confirms each finding, and false-positive rate is measured, not asserted.

Raw AI output passed through with no human triage, which is exactly the source of the 78 percent false-negative problem above.

3. If you run an autonomous agent against our systems, what runtime controls and scope enforcement are in place?

Enforced scope, rate limits, an immediate stop control, and a full audit trail of every action the agent took.

An autonomous tool pointed at production with no scope enforcement or kill switch.

4. How is our data used with any AI tooling: is it sent to third-party models, and is it used for training?

Clear data-flow disclosure, no training on your data, and named model or self-hosted arrangement with contractual backing.

Vague "we use AI" with no answer on where your data goes or whether it trains a third-party model.

5. What can your AI find that a scanner cannot, and where does it still need a human?

Honest scope: strong on coverage and speed, with humans required for business logic, chained exploits, and impact validation.

A claim that AI fully replaces human testers, which no current benchmark supports.

Remaining AI and automation questions to include:

  1. What is the measured false-positive and false-negative rate of your automated tooling, and how do you measure it?

  2. How do you keep AI tooling from acting outside the agreed scope during an unattended run?

  3. Do you offer a hybrid model, and how is human review structured within it?

  4. How do you handle prompt-injection and data-exfiltration risk in your own AI tooling?

  5. Can you run a small proof of concept on our own application before we commit?

  6. How do you version and log AI tool behavior so a finding is reproducible months later?

  7. What happens to AI-generated evidence and logs at the end of the engagement?

For a deeper vendor-side interrogation of AI claims, use the questions to ask an AI pentest vendor, and to prove the claims on your own app, run the 30-day AI pentest bake-off scorecard.

Section 5: Reporting quality (weight 15)

The report is the deliverable you actually keep. A strong report is reproducible, prioritized by real risk, and usable by both an engineer fixing the bug and an auditor reviewing the evidence. Ask for a sample before you sign, because a redacted real report tells you more than any proposal paragraph.

Question

A strong answer contains

Red-flag answer

1. Can you provide a redacted sample report from a comparable engagement?

A real, redacted report with clear reproduction steps, evidence, severity justification, and remediation guidance.

Refusal to share any sample, or a one-page "certificate" with no technical detail.

2. What does a single finding contain, from reproduction steps to remediation?

Reproduction steps, proof-of-concept evidence, business impact, CVSS or rated severity, and specific, testable remediation.

Findings that are just scanner output with a generic "update to latest version" fix.

3. How do you tie findings to standards and compliance evidence we can hand to an auditor?

Mapping to the frameworks you name (SOC 2, ISO 27001, PCI DSS, DORA), so the report doubles as audit evidence.

No mapping, leaving you to translate raw findings into audit language yourself.

4. What do you deliver beyond the PDF: readout, developer walkthrough, machine-readable output?

A findings review call, a developer-facing walkthrough, and optional structured output for your ticketing system.

Report emailed with no debrief and no path to the people who have to fix the issues.

Remaining reporting questions to include:

  1. How do you prioritize findings when everything cannot be fixed at once?

  2. How do you handle false positives found after delivery, and do you correct the report?

  3. What is your turnaround from end of testing to draft report, and to final report?

  4. Do you provide an executive summary that a non-technical stakeholder can act on?

  5. How do you present findings that are chained, where two medium issues combine into a critical?

  6. Will you present to our board or auditors if we need you to, and at what cost?

For a full rubric on grading the deliverable itself, see how to evaluate a penetration test report.

Section 6: Retesting and remediation (weight 10)

A finding you cannot prove is fixed is a finding you cannot close. Retesting is where thin bids quietly save money, by excluding fix verification, so the true cost lands on you the first time an auditor asks for evidence of remediation. Ask exactly what retesting is included and what triggers a paid re-scope.

Question

A strong answer contains

Red-flag answer

1. Is retesting of fixed findings included, and for how long after delivery?

Retesting included within a defined window (commonly 30 to 90 days), with a clear statement of what counts as a fix verification versus a new test.

Retesting billed as a full new engagement, or no retest offered at all.

2. How do you verify a fix: targeted re-test of the finding, or a full re-scan?

A targeted re-test that confirms the specific issue is resolved without re-charging for the whole engagement.

Only a full re-scan, which inflates cost and slows down your remediation cycle.

3. How do you support our developers during remediation, and is that time included?

Access to the tester for remediation questions, and clear guidance on fixes, within the engagement scope.

Remediation support treated entirely as billable extra with no included support at all.

Remaining retesting and remediation questions to include:

  1. What is your SLA for confirming a critical-severity fix?

  2. How do you track remediation status across multiple findings over time?

  3. Do you offer continuous or periodic retesting, and how is that priced against a one-off test?

  4. How do you handle a finding we dispute or accept as a risk rather than fix?

  5. What happens if a fix introduces a new vulnerability that your retest uncovers?

For the contractual side of retest guarantees and turnaround, see autonomous pentest contract and SLA clauses.

This section protects you from turning a security test into a data-protection incident. You are granting a third party access to your systems and, often, your data, so the RFP has to pin down where that data lives, how long it is kept, and who is liable if the engagement goes wrong.

Question

A strong answer contains

Red-flag answer

1. Where is our data stored during and after the engagement, and for how long is it retained?

Named storage locations, a defined retention period, and a documented secure-destruction step with confirmation.

No clear answer on where data lives, or indefinite retention with no destruction policy.

2. What contractual protections do you offer: NDA, liability cover, data-processing terms?

A signed NDA, professional indemnity insurance, and data-processing terms that match your regulatory obligations.

Reluctance to sign an NDA, or no liability cover and no data-processing agreement.

3. How do you ensure testing does not disrupt production or expose real customer data?

Non-destructive testing rules, staging where appropriate, and explicit handling for any real data encountered.

No safeguards for production, or a willingness to test destructively without written authorization.

Remaining data handling and legal questions to include:

  1. How do you securely transmit findings and evidence to us?

  2. Who on your side has access to our data, and how is that access controlled and logged?

  3. How do you handle a breach of your own systems that could expose our engagement data?

  4. What is your subprocessor list, and how are subprocessors bound to the same terms?

  5. How do you support our own regulatory obligations (GDPR, sector rules) as a data processor?

A worked set of strong answers

The fastest way to calibrate the red-flag key is to read a set of strong answers end to end. As the buyer, you are entitled to hold your shortlist to this standard. Here is how a hybrid, human-plus-AI provider answers a handful of the hardest questions in this bank, written the way you should expect a serious bidder to answer.

On firm accreditation and tester qualifications (Section 2, Q1 and Q4). Stingrai is a CREST-accredited firm, and its engagements are delivered by senior human pentesters whose certifications include OSCE3, OSCP, OSWE, CREST CRT, and CRTO, alongside 18 published CVEs and research presented at DEF CON and BSides. Testers are named before the engagement, and for regulated testing the qualification mapping is explicit rather than a blanket compliance claim.

On methodology and manual depth (Section 1, Q1 to Q3). The work runs on a named methodology, and reporting is mapped to MITRE ATT&CK so red team coverage is legible rather than asserted. Manual effort concentrates on the classes automated tools miss: broken authorization, IDOR, and business-logic flaws.

On AI disclosure and the hybrid model (Section 4, Q1 and Q5). Stingrai delivers a transparent hybrid model. Snipe, its autonomous web-application testing agent, is purpose-built to hunt the complex classes generic scanners miss, IDOR, broken access control, and business-logic flaws, and it works black-box and white-box, reviewing source code, opening AutoFix pull requests, and running as a pull-request gate that can block vulnerable code from merging. It is trained on more than 6,000 HackerOne Hacktivity disclosure reports and on skills distilled from the firm's own human pentesters. Cloud, identity, network, and social-engineering testing is led by human pentesters. Every high-impact finding is validated by a senior human before it reaches your report, which is the step that keeps false positives off the page.

On retesting and evidence (Section 5 and Section 6). Retesting and AutoFix pull requests are part of the delivery model rather than a billable afterthought, and the reporting is built to support your SOC 2, ISO 27001, PCI DSS, and DORA evidence needs so the deliverable does double duty as audit material. Pricing for these packages is published openly at stingrai.io/pricing.

You do not have to shortlist any particular vendor. You do have to hold whoever you shortlist to answers at this level of specificity. If a bidder cannot, the red-flag key will tell you in a single scoring pass. To see the full service scope behind these answers, review web application penetration testing, red teaming, and the wider penetration testing services.

Frequently asked questions

What questions should I include in a penetration testing, red team, or cloud pentest RFP, and what vendor answers are red flags?

Include questions across seven areas: general methodology, tester qualifications and red team capability, cloud and infrastructure depth, AI and automation disclosure, reporting quality, retesting, and data handling. The clearest red flags are a scanner-only workflow sold as a manual pentest, testers the vendor will not name or credential, no sample report, findings with no retest included, an autonomous AI tool run against production with no scope enforcement, and any refusal to say which findings a human validated. This bank scores all seven areas on a 0 to 5 scale so each red flag becomes a number.

What is a penetration testing RFP template?

A penetration testing RFP template is a structured request for proposal that a buyer sends to security testing vendors to compare them on the same criteria. A strong template goes beyond a question list: it weights the sections, scores each answer, and defines the red-flag answers in advance, so replies from different vendors are genuinely comparable instead of a stack of confident prose. This bank is a scored template you can copy directly into your procurement process.

How do you score penetration testing vendor responses in an RFP?

Weight the sections to your own risk so they total 100, then score each answer from 0 to 5, where 0 is a documented red-flag answer, 3 is adequate, and 5 is the strong answer defined in the key. Multiply each section's average by its weight and total to 100. Add an override rule: a red-flag answer on any high-weight question caps that section regardless of the rest of the reply, which stops a strong sales narrative from hiding a fundamental gap.

What are red flags in a penetration testing vendor's proposal?

The most common red flags are automation sold as manual testing, unnamed or uncredentialed testers, no redacted sample report, business logic and authorization left out of scope, findings delivered with no retest, and vague answers about where your data goes and how long it is kept. In AI-assisted testing, the added red flags are raw automated output with no human validation and any autonomous agent pointed at production without enforced scope and a stop control.

What red team tester qualifications should a DORA TLPT or TIBER-EU RFP require?

Lift the thresholds straight from the rules. The DORA threat-led penetration testing Regulatory Technical Standards (EU 2025/1190, Article 7) require a lead tester with at least five years of penetration and red team testing experience, at least two more testers with at least two years each, a combined five or more prior assignments, at least five references, CVs with market-standard certifications, and professional indemnity insurance. The ECB's TIBER-EU guidance adds a red team manager with at least five years including three years leading financial-sector tests, and independence between the threat-intelligence and red team functions.

What AI and automation disclosure questions should a pentest RFP include?

Ask whether the vendor uses AI or autonomous agents, which findings come from automation versus a human, how false positives are prevented, what runtime scope controls and stop conditions apply to any autonomous agent, and whether your data is sent to third-party models or used for training. This matters because support for fully automated pentesting fell from 29 percent to 9 percent in a year and 78 percent of teams hit critical false negatives from automated tools (Cobalt, 2026). Disclosure and human validation, not the presence of AI, are what you are scoring.

What cloud pentest RFP criteria matter most?

The criteria that separate a real cloud test from a network scan are identity and access management testing (roles, trust policies, privilege escalation), control-plane and configuration review against a named benchmark, container and Kubernetes security, and a least-privilege access model for any credentials you issue. A vendor that answers "we scan the external IP range" for a cloud environment is testing the wrong layer, because most cloud compromise starts in identity and configuration, not the network edge.

What should a pentest RFP ask about reporting quality?

Ask for a redacted sample report before signing, and confirm that each finding contains reproduction steps, evidence, business impact, a rated severity, and specific remediation. Ask how findings map to the compliance frameworks you care about, and what you get beyond the PDF: a findings review call, a developer walkthrough, and machine-readable output. A one-page certificate with no technical detail is a red flag, because the report is the deliverable you keep and hand to auditors.

What should a pentest RFP ask about retesting?

Ask whether retesting of fixed findings is included and for how long, whether a fix is verified with a targeted re-test or a full re-scan, and what SLA applies to confirming a critical fix. Retesting is where thin bids hide cost by excluding fix verification, so pin it down in the RFP. A vendor that only offers a full new engagement to confirm a fix will slow your remediation and inflate your spend.

What questions should a pentest RFP include about data handling?

Ask where your data is stored during and after the engagement, how long it is retained, and how it is securely destroyed. Confirm the vendor will sign an NDA, carries professional indemnity insurance, and offers data-processing terms that match your regulatory obligations. Ask who on their side can access your data, how that access is logged, and what their subprocessor list looks like. Vague retention answers or reluctance to sign an NDA are red flags.

References

  1. Cobalt. AI and Pentesting Pulse Report 2026. June 2026. https://resource.cobalt.io/ai-pentesting-pulse-report-2026-tyd. Survey of 455 security leaders and practitioners built on five years of pentesting data; source for the 29 percent to 9 percent drop in support for fully automated pentesting, the 78 percent critical-false-negative figure, and the 47 percent hybrid-model preference.

  2. HackerOne. 2025 Hacker-Powered Security Report. November 2025. https://www.hackerone.com/blog/ai-security-trends-2025. Source for the 210 percent year-over-year jump in valid AI-assisted vulnerability reports.

  3. European Commission. Commission Delegated Regulation (EU) 2025/1190 (DORA TLPT Regulatory Technical Standards). Applicable 8 July 2025. https://eur-lex.europa.eu/eli/reg_del/2025/1190/oj. Source for threat-led penetration testing tester-qualification requirements in Article 7: lead experience, additional testers, references, certifications, and insurance.

  4. European Central Bank. TIBER-EU Framework and Service Provider Procurement Guidance. https://www.ecb.europa.eu/paym/cyber-resilience/tiber-eu/html/index.en.html. Source for red team manager and red team member experience thresholds and independence between threat intelligence and red team roles.

  5. Bank of England. CBEST Threat Intelligence-Led Assessments Implementation Guide. https://www.bankofengland.co.uk/financial-stability/operational-resilience-of-the-financial-sector/cbest-threat-intelligence-led-assessments-implementation-guide. Source for CREST accreditation and certification expectations governing intelligence-led financial-sector testing.

  6. CREST. Certifications and accreditation. https://www.crest-approved.org/. Source for the red team and threat-intelligence certifications referenced in the tester-qualification section.

Score your shortlist against real answers

You now have a scored RFP question bank, a red-flag answer key, and the regulator-grade tester thresholds to hold vendors to. The last step is comparing the replies against a provider that answers at this level of specificity. Stingrai, a CREST-accredited firm, delivers hybrid penetration testing and red teaming that pairs senior human testers with Snipe, its autonomous web-application agent, with named methodology, MITRE ATT&CK-mapped reporting, AutoFix pull requests, included retesting, and evidence built to support your SOC 2, ISO 27001, PCI DSS, and DORA programs. See the penetration testing services, the PTaaS platform, and open pricing to benchmark your shortlist.

0 views

0

X

Related reading

How to Scope a SaaS OAuth and Connected App Penetration Test
Web App SecurityNetwork Security

How to Scope a SaaS OAuth and Connected App Penetration Test

Scope a SaaS OAuth and connected app penetration test: connected app inventory, token scope review, consent grant hygiene, and blast radius testing.

11 min read

AEV, BAS, PTaaS, or Autonomous Pentest? A Buyer's Decoder for Gartner's New Categories
Network SecurityWeb App Security

AEV, BAS, PTaaS, or Autonomous Pentest? A Buyer's Decoder for Gartner's New Categories

Gartner's 2026 AEV category folds in BAS and automated pentesting. This decoder maps AEV, BAS, CTEM, PTaaS and autonomous pentest to what each one proves.

16 min read

How to Scope a Google Cloud Penetration Test: The GCP-Native Surface an AWS Guide Misses
Network SecurityWeb App Security

How to Scope a Google Cloud Penetration Test: The GCP-Native Surface an AWS Guide Misses

Scope a Google Cloud penetration test the GCP-native way: resource hierarchy, service accounts, IAM Conditions, GKE, and Workspace, not an AWS checklist.

11 min read

Contents

X