A penetration testing RFP is only as strong as the answers it lets you compare. Most templates hand you a list of questions and no way to grade the replies, so every vendor reads as competent and the cheapest bid wins by default. This bank is built the other way around. It gives you 75 vendor questions across seven weighted sections that total 100, a scored answer key that spells out what a strong reply contains and what a red-flag reply sounds like, a dedicated AI and automation disclosure section, and tester-qualification questions drawn straight from the regulated red team frameworks. Those four things, weighted scoring, a red-flag key, AI disclosure, and regulator-grade tester questions, are exactly what a generic template leaves out.
It also gives you the document itself. The copy-ready RFP template further down supplies nine parts you can paste straight into your own file: company background and firm qualification, a scope definition matrix, methodology and standards requirements, tester qualifications with a procedure for verifying a firm-level CREST accreditation, evidence and reporting requirements, service levels and retest terms, a commercial response format, the evaluation scoring rubric you publish to bidders, and the automatic disqualifiers that end a bid. Nothing on this page is gated and nothing asks for your email.
The core question this answers: what questions should you include in a penetration testing, red team, or cloud pentest RFP, and which vendor answers are red flags? Ask about methodology, tester qualifications, cloud depth, AI and automation disclosure, reporting quality, retesting, and data handling. Treat these replies as red flags: a scanner-only workflow sold as a pentest, testers the vendor will not name or credential, no sample report, findings with no retest, and any refusal to say which findings a human actually validated. Every one of those is scored below, so the flag turns into a number rather than a gut feeling.
This is written for the procurement lead, security engineer, or CISO who is drafting a vendor questionnaire right now and needs the replies to be comparable when they come back. Copy the template, paste the questions into it, set the weights to your own risk, and hand your evaluators the answer key so three people scoring the same proposal land within a point of each other. If you are earlier in the process and still comparing prices, start with how to compare penetration testing quotes, then come back here to score the technical response. If you are still building the longlist, penetration testing vendors covers who to invite before you score anyone.
Key takeaways
The template is copy-ready, not a description of one. Nine parts, built as tables and checklists, cover everything from the scope definition matrix to the automatic disqualifiers. Paste them into your document, replace the bracketed placeholders, and issue it.
Score the answers, do not just collect them. A template that gathers replies without a rubric lets every vendor sound equal. Weighting seven sections to 100 and scoring each answer 0 to 5 turns a stack of proposals into a ranking your whole committee can defend.
AI disclosure is the new must-ask section. Support for fully automated pentesting fell from 29 percent to 9 percent in a single year, and 78 percent of teams reported fully automated tools missing critical vulnerabilities and returning false negatives (Cobalt, 2026). An RFP that does not ask what runs the test is buying blind.
Regulated red team work already sets a tester-qualification bar. The DORA threat-led penetration testing rules require a lead with at least five years of penetration and red team experience plus two more testers with at least two years each (EU 2025/1190, Article 7). You can lift that bar into any red team RFP as pass-or-fail criteria.
Banks and insurers have a second layer of criteria. In regulated sectors the tester bar, the independence rules, and the testing cadence are set by the supervisor, not by preference. The regulated-sector overlay below lifts the exact thresholds out of DORA TLPT, TIBER-EU, CBEST, APRA CPS 234, and the MAS Technology Risk Management Guidelines and turns them into pass-or-fail prequalification.
A scanner sold as a pentest is the most common red flag. The cheapest way to lose the value of a test is to buy an automated scan dressed up as manual testing. Every methodology and reporting question below is designed to surface that substitution before you sign.
Retesting and data handling are where thin bids hide cost. A proposal that excludes fix verification, or that is vague about where your data lives and how long it is kept, is cheaper today and more expensive the first time you need evidence for an audit or an incident.
Methodology and sources
This bank is a practitioner tool, not a survey, but the numbers and thresholds inside it are sourced to named primary publishers so you can cite them in your own RFP:
Cobalt, AI and Pentesting Pulse Report 2026 (June 2026): a survey of 455 security leaders and practitioners built on five years of pentesting data. Source for the year-over-year drop in support for fully automated pentesting (29 percent to 9 percent), the 78 percent critical-false-negative figure, and the 47 percent hybrid-model preference.
HackerOne, 2025 Hacker-Powered Security Report (November 2025): source for the 210 percent year-over-year jump in valid AI-assisted vulnerability reports.
DORA TLPT Regulatory Technical Standards, Commission Delegated Regulation (EU) 2025/1190 (applicable 8 July 2025): source for the threat-led penetration testing tester-qualification requirements in Article 7.
TIBER-EU framework and service provider procurement guidance, European Central Bank: source for the red team manager and red team member experience thresholds.
CBEST implementation guide, Bank of England, and CREST: source for the CREST accreditation and certification expectations that govern intelligence-led financial-sector testing.
APRA Prudential Standard CPS 234, Information Security (Australian Prudential Regulation Authority): source for the systematic testing program requirement, the requirement that testing be conducted by appropriately skilled and functionally independent specialists, and the third-party testing and notification obligations quoted in the regulated-sector overlay.
MAS Technology Risk Management Guidelines (Monetary Authority of Singapore, January 2021): source for the penetration testing scope and frequency expectations and the adversarial attack simulation exercise requirements quoted for APAC financial institutions.
The research cutoff for this pass was July 2026. Where a figure could not be confirmed against its named primary publisher on this pass, it was left out rather than estimated. Every stat in the sections below links back to its publisher so any claim can be audited inline.
What to put in an RFP, by buying scenario
Three buying scenarios account for most of the questions procurement teams bring to this document. Each one shifts which sections carry the weight and adds criteria the others do not need, so start with the scenario that matches your purchase, then use the full bank and the copy-ready template that follow.
What criteria should be included in an RFP for an enterprise autonomous pentest?
An RFP for an enterprise autonomous pentest should score six criteria a traditional pentest RFP never had to ask about: runtime scope enforcement with a working stop control, human validation of every high-severity finding before it reaches your report, measured false-positive and false-negative rates, disclosure of where your data goes and whether it trains a third-party model, a reproducible audit trail of every action the agent took, and evidence of which vulnerability classes the agent actually reaches. Then add the commercial criteria that autonomous delivery changes: pricing per asset or per subscription rather than per engagement day, how assets are onboarded and removed mid-term, and how findings land in your existing engineering workflow.
Score these criteria, in this order:
Scope enforcement and stop control. Enforced target boundaries, rate limits, an immediate kill switch, and a named person on both sides who can invoke it.
Human validation. A senior tester confirms every high-severity finding before delivery, and the vendor can show you which findings a human touched.
Measured accuracy. A stated false-positive and false-negative rate with the measurement method behind it, not an adjective.
Vulnerability class coverage. Evidence the agent reaches broken authorization, IDOR, and business-logic flaws, not only the known classes a scanner already finds.
Data handling. Whether your code and traffic leave your boundary, which models see them, and whether they are used for training.
Reproducibility. A full action log and versioned tool behavior, so a finding can be reproduced months later during an audit or a dispute.
Workflow integration. Findings pushed into your ticketing system, and optionally pull-request gating and automated fix suggestions in your repository.
Commercial model. Per-asset or subscription pricing, mid-term changes to the estate, and a defined exit with data destruction.
Proof before commitment. A proof of concept against one of your own applications, scored against a finding you already know about.
Section 4 below turns each of these into scored questions with a red-flag key, and the questions to ask an AI pentest vendor goes further on interrogating the claims themselves.
What should a bank put into an RFP for red team services?
A bank procuring red team services should build the RFP around four things a general pentest RFP leaves out: the supervisor's tester-qualification thresholds written in as pass-or-fail prequalification, enforced independence between the threat-intelligence function and the red team, a rules-of-engagement and deconfliction regime that protects production and customer data, and objectives defined against critical business functions rather than a list of IP ranges. Then add measurement of your own defenders, because the deliverable that matters to a regulated firm is evidence about detection and response, not a vulnerability list.
Lift these into the document:
Prequalification thresholds from the applicable regime. DORA TLPT and TIBER-EU in the EU, CBEST in the UK, CPS 234 in Australia, and the MAS adversarial attack simulation expectations in Singapore. The regulated-sector overlay below gives the exact thresholds with their primary sources.
Threat-intelligence independence. A separate provider or a walled-off internal team, so the scenarios reflect adversaries that actually target institutions like yours.
Crown-jewel objectives. Named critical business functions and the flags that prove reach, agreed in writing before the test starts.
Rules of engagement, deconfliction, and stop authority. Signed before any activity, with named contacts on both sides and an immediate halt that does not route through an account manager.
Production safeguards. Written constraints covering destructive actions, customer data, and payment flows.
Defender measurement. Detection and response timings captured per scenario, plus a purple-team replay so your team learns from what was missed.
Board-grade reporting. Findings your board and your supervisor can read, with any control deficiency that cannot be remediated quickly called out explicitly.
Data residency, retention, and insurance. Where engagement data lives, how long it is kept, and the indemnity cover standing behind it.
Sections 2 and 7 below score most of this. For the regional shortlist, the country rankings cover the UK, Australia, Singapore, and Canada.
How do you put together an RFP for continuous penetration testing?
An RFP for continuous penetration testing has to define what continuous actually means before it asks anyone for a price, because the word currently covers three different products: always-on automated scanning, scheduled waves of human testing spread across the year, and testing gated on every code change. State which one you are buying, then specify the asset inventory and the change triggers that pull an asset back into testing. The commercial section carries more weight than in a one-off RFP because you are signing a term rather than a project, so pin down per-asset pricing, mid-term onboarding and removal, retest turnaround, and exit before you compare totals.
Specify these in the document:
Definition. Which of the three models you are buying, and the human testing hours included per period.
Asset inventory and change triggers. The applications, APIs, and environments in scope, plus the events that force a new test: a major release, an architecture change, a new authentication flow, a new third-party integration.
Cadence and coverage. How often each asset is tested, and how the vendor proves coverage instead of repeatedly retesting the easy surface.
Retest service level. Turnaround to verify a fix, included in the subscription rather than billed on top.
Delivery into your workflow. Findings pushed to your ticketing system with severity and reproduction steps, not a quarterly PDF that nobody opens.
Periodic summary reporting. A scheduled report your auditors can use as evidence that testing actually ran across the term.
Commercial terms. Per-asset or subscription pricing, what happens when the estate grows, notice period, and exit with confirmed data destruction.
The contract-side detail behind all of this, including how to write retest and turnaround guarantees that survive a continuous or autonomous delivery model, is in autonomous pentest contract and SLA clauses.
How to use this bank

The seven sections carry a default weight that sums to 100. Adjust the weights to your own risk, a cloud-native company might push cloud depth to 20 and trim red team, a regulated bank will do the reverse, but set the weights before you read a single proposal and keep them identical across every bidder. Changing weights mid-evaluation is how a committee talks itself into the vendor it already liked.
Score each answer on a 0 to 5 scale: 0 for a red-flag answer, 3 for adequate, 5 for the strong answer described in the key. Multiply each section's average by its weight, then total to 100. Add one override rule that protects the whole exercise: a red-flag answer on any high-weight question caps that section, no matter how polished the rest of the reply reads. A vendor that cannot name its testers does not get to make it up on price.
The strongest 30 questions below ship with a full answer key in table form. The remaining questions are listed compactly by section so you can drop the whole bank into your RFP and still keep the document readable. The copy-ready template further down is where these seven sections go: it is the surrounding document, and the questionnaire is Part 3 and Part 4 of it. If you want a scope document to sit alongside the questionnaire, the AI and LLM pentest scope-of-work template pairs cleanly with this bank.
The red-flag answer key, explained

Each scored question has three parts: the question you ask, the signals a strong answer contains, and the red-flag answer that caps or fails the score. The point of writing the red flag down in advance is consistency. When three evaluators independently know that "we run an automated scanner and pass you the output" is a 0 on a manual-methodology question, they score the same proposal the same way. Without the key, one evaluator hears "we use automation for efficiency" as innovative and another hears it as a substitution, and your ranking becomes an argument about tone instead of substance.
The copy-ready RFP template
Everything in this section is written to be lifted. Copy the headings, tables, and checklists into your own document, replace the bracketed placeholders with your details, and delete the parts that do not apply. The nine parts below are the document skeleton. The seven scored sections that follow are what you paste into Part 3 and Part 4 as the technical questionnaire.
Part 1: Company background and firm qualification
State what the bidder must supply about itself, and state what you will independently verify. Verification matters more than the claim, because accreditations and insurance are checkable, and a bidder that resists verification has already told you something useful.
Requirement | What the bidder must supply | How we verify it |
|---|---|---|
Legal entity and trading history | Registered legal name, country of incorporation, and years delivering this specific service | Company register, and the entity named on the accreditation must match the entity that signs the contract |
Firm-level accreditation | Accrediting body, the service scope of the accreditation, and current status | The accrediting body's public member directory, checked by us |
Individual certifications | Named testers with certification names and numbers | Certifying body verification, plus CVs supplied with the bid |
Relevant sector experience | Engagements of comparable scope, sector, and technology in the last 24 months | Two client references we may contact directly |
Delivery locations and subcontracting | Where the testers sit, and any subcontracted work disclosed up front | Written declaration, binding for the term |
Insurance | Professional indemnity and cyber liability cover, with limits stated | Certificate of insurance naming the bidding entity |
Conflicts of interest | Any product the bidder resells or implements that appears in our environment | Written declaration |
Part 2: Scope definition matrix
The single biggest cause of non-comparable bids is a scope description loose enough that each vendor assumes something different. Fill this matrix in before you issue the RFP, and require every bidder to price against it exactly.
Asset or system | Type | Environment | Size or count | Test window | Constraints |
|---|---|---|---|---|---|
[Customer web application] | Web application | [Staging with production-like data] | [User roles, key flows] | [Dates] | [No destructive testing, no load testing] |
[Public API] | REST or GraphQL API | [Production] | [Endpoints, authentication model] | [Dates] | [Rate limit ceiling] |
[External perimeter] | Network | Production | [Live hosts or IP ranges] | [Dates] | [Change freeze dates] |
[Internal network] | Network | Production | [Subnets, domain size] | [Dates] | [Assumed-breach starting position] |
[Cloud environment] | AWS, Azure, or GCP | Production | [Accounts or subscriptions] | [Dates] | [Read-only roles issued, no data exfiltration] |
[Kubernetes estate] | Container platform | [Environment] | [Clusters and namespaces] | [Dates] | [No cluster-level disruption] |
[Mobile application] | iOS, Android | [Build channel] | [Builds] | [Dates] | [Test accounts supplied by us] |
[Social engineering] | Phishing or vishing | Production | [Target population] | [Dates] | [No credential capture beyond first factor] |
[Red team objective] | Adversary emulation | Production | [Named critical business function] | [Dates] | [Deconfliction contacts, stop authority] |
Publish these scope facts with the RFP so no bidder has to guess and then hedge the price: the number of applications and distinct user roles, the authentication and authorization models, endpoint and IP counts, whether testing runs against production, any change freeze windows, the credentials and test accounts you will issue, and the data classification of anything a tester might encounter.
Part 3: Methodology and standards requirements
Require each bidder to answer these as statements about their own process, not as a brochure. Each one is scored using Section 1 below.
Name the methodology and standards applied, and explain how they are tailored to our architecture.
State the manual versus automated split, and label which findings came from which.
Describe the manual test cases used for broken authorization, IDOR, and business-logic flaws.
State the severity model, and how business impact is assessed separately from raw tool output.
For red team work, map the planned activity to MITRE ATT&CK techniques.
Describe the evidence standard: what proof accompanies a finding.
Describe the internal quality process that stops a weak test shipping as a confident report.
State the change-control process for scope discovered mid-test.
Part 4: Tester qualifications, and how to verify a CREST accreditation
Set the minimums as pass-or-fail, then verify them instead of accepting the claim.
Requirement | Minimum we require | Evidence required with the bid |
|---|---|---|
Lead tester experience | [5] years hands-on in this discipline | CV with dates, plus certification numbers |
Additional testers | [2] testers with at least [2] years each | CVs with dates |
Certifications | At least one of OSCP, OSCE3, OSWE, CREST CRT, or CRTO per assigned tester | Certification numbers we can verify with the issuing body |
Firm accreditation | Current firm-level accreditation covering the service being bought | Entry in the accrediting body's directory |
Comparable assignments | [5] engagements of comparable scope in the last 24 months | Anonymized summaries plus [2] contactable references |
Named delivery | The named testers are the people who actually deliver | Contractual commitment, with substitution requiring our written approval |
Firm-level accreditation is the requirement most often misread in a bid, because a vendor can truthfully say the word "CREST" when what it means is that one employee once passed a CREST exam. Those are different things, and only one of them tells you anything about the company. Verify it in five steps:
Ask for the exact legal entity name that holds the accreditation, then check it matches the entity that will sign your contract. Group structures often hold the accreditation in one subsidiary and deliver the work from another.
Search the accrediting body's public member directory rather than trusting a logo on a website. CREST publishes its member companies at marketplace.crest.org.
Confirm which service the accreditation covers. Accreditation is granted per discipline, so a firm accredited for one service is not automatically accredited for penetration testing, red teaming, or threat intelligence. Ask which discipline, and check it against what you are actually buying.
Separate firm accreditation from individual certification. Require both and evidence them separately: the company's accreditation on one line, the named individuals' certification numbers on another.
Ask for current status and renewal date, and require the bidder to notify you in writing if the accreditation lapses during the term.
For the wider list of accredited providers and what the accreditation actually covers, see CREST-accredited penetration testing companies. For how firms compare on capability before you invite them, see penetration testing vendors.
Part 5: Evidence and reporting requirements
State the deliverable in the contract, so the report you receive is the report you scoped.
A redacted sample report from a comparable engagement, supplied with the bid rather than after award.
Every finding contains: title, affected asset, reproduction steps, evidence, business impact, rated severity, and specific remediation guidance.
Chained findings are presented as chains, with the combined impact rated rather than the parts rated separately.
An executive summary a non-technical stakeholder can act on without translation.
Mapping of findings to the frameworks we name: [SOC 2, ISO 27001, PCI DSS 4.0, DORA, NIS2].
Machine-readable output we can import into our ticketing system.
A findings review call and a developer-facing walkthrough, both included.
A retest report stating, finding by finding, whether the fix was verified.
A summary letter confirming the engagement took place, for our own customers who ask for one.
Named format, delivery mechanism, and encryption for every deliverable.
Write the framework mapping requirement carefully, because the frameworks expect different things from scope, frequency, and evidence, and a generic "maps to compliance" line lets a bidder satisfy none of them properly. Before you fill in that bracket, check what your framework actually asks for: SOC 2 penetration testing and PCI DSS penetration testing cover the scope, cadence, and evidence expectations for each.
Part 6: Service levels, retesting, and contract terms
Term | Our requirement | Bidder response |
|---|---|---|
Critical finding notification | Notified within [24] hours of discovery, ahead of the report | |
Draft report turnaround | Within [5] business days of test completion | |
Final report turnaround | Within [5] business days of our comments | |
Retest window | Included for [90] days after the final report | |
Retest turnaround | Within [10] business days of our fix notification | |
Remediation support | [8] hours of tester time included for fix questions | |
Data retention and destruction | Retained no longer than [90] days, destroyed with written confirmation | |
Data location | Engagement data stored in [jurisdiction] only | |
Subprocessors | Disclosed list, bound to the same terms, changes notified [30] days ahead | |
Confidentiality | Signed NDA before any scoping detail is exchanged | |
Liability and insurance | Professional indemnity of at least [amount], cyber liability of at least [amount] | |
Stop-work authority | We may halt testing immediately by phone or in writing, via named contacts | |
Right to audit | We may review how the bidder handles our engagement data | |
Term and exit | [12] months, [60] days notice, data destroyed at exit |
The clause-level detail behind these terms, including how to write retest and turnaround guarantees that survive an autonomous or continuous delivery model, is in autonomous pentest contract and SLA clauses.
Part 7: Commercial response format
Require one format so bids are comparable line by line. A bidder that responds with a single blended number has made itself impossible to score, which is often the intention.
Line item | Unit | Quantity | Unit price | Total |
|---|---|---|---|---|
Web application test | Per application | |||
API test | Per API | |||
External network test | Per IP range | |||
Internal network test | Per engagement | |||
Cloud environment review | Per account or subscription | |||
Red team engagement | Per engagement | |||
Social engineering campaign | Per campaign | |||
Retesting | Included, or per finding | |||
Remediation support | Per hour | |||
Continuous or subscription option | Per asset per year | |||
Travel and expenses | At cost, capped |
Also require, in writing: what is explicitly excluded, what triggers a change order, the payment milestones, and the price validity period. For normalizing the answers that come back, use how to compare penetration testing quotes.
Part 8: Evaluation and scoring rubric
Publish this to bidders with the RFP. Telling vendors how you score raises the quality of what you receive and removes the argument afterwards.
Evaluation area | Weight | Scored using |
|---|---|---|
General and methodology | 15 | Section 1 |
Tester qualifications and red team | 20 | Section 2 |
Cloud and infrastructure depth | 15 | Section 3 |
AI and automation disclosure | 15 | Section 4 |
Reporting quality | 15 | Section 5 |
Retesting and remediation | 10 | Section 6 |
Data handling and legal | 10 | Section 7 |
Total | 100 |
State the mechanics in the document itself: each answer scores 0 to 5, section scores are averaged and multiplied by the section weight, and a red-flag answer on any high-weight question caps that section regardless of how the rest of the reply reads. Say whether price is scored inside this 100 or evaluated separately against the shortlist, and say it before proposals arrive rather than after you have read them.
Part 9: Automatic disqualifiers
List these in the RFP so bidders self-select out before they waste your evaluation time.
Automated scanning presented as manual penetration testing.
Refusal to name the testers who will deliver, or to supply their CVs.
Refusal to supply a redacted sample report.
No retesting of fixed findings included in any form.
Undisclosed subcontracting of the testing work.
An autonomous testing agent with no scope enforcement and no stop control.
No answer on where engagement data is stored or how long it is retained.
Unwillingness to sign an NDA or to provide proof of insurance.
A firm-level accreditation claim we cannot find in the accrediting body's directory.
Submission instructions and timeline
Milestone | Date |
|---|---|
RFP issued | [date] |
Bidder questions due | [date] |
Answers published to all bidders | [date] |
Proposals due | [date] |
Shortlist notified | [date] |
Shortlist presentations and technical deep dive | [date] |
Award decision | [date] |
Engagement kickoff | [date] |
Include the submission address, the required format and page limit, the single named contact who may be approached during the process, the proposal validity period, and a statement that answers to bidder questions will be shared with every bidder.
Nothing on this page is gated, so take the template and issue it with whatever vendor list you like. If you would rather have a provider respond to the finished document, send your RFP to Stingrai or request a quote.
Section 1: General and methodology (weight 15)
This section separates a structured, standards-based engagement from an ad hoc scan. You are looking for a named methodology, manual testing of the classes automated tools miss, and a scope conversation that starts with your risk rather than a fixed price list.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. What methodology and standards do you follow, and how do you adapt them per engagement? | Named framework (OWASP WSTG, PTES, NIST SP 800-115, OSSTMM) applied as a baseline, then tailored to the target's architecture and threat model with examples. | "We use our own proprietary process" with no named standard and no detail, or a pure tool list. |
2. What share of the test is manual versus automated, and which findings come from which? | A clear split, with manual effort focused on authorization, business logic, and chained exploits, and automation limited to coverage and enumeration. | "It is fully automated" or an inability to say which findings a human produced. |
3. How do you scope authorization, IDOR, and business-logic flaws that a scanner cannot see? | Concrete manual test cases per role and object, multi-account testing, and worked examples from a sample report. | Business logic treated as out of scope, or folded into a scanner run with no manual cases. |
4. How do you define and justify severity, and do you use a named rating system? | CVSS or a documented internal model tied to real business impact and exploitability, not raw scanner severities. | Severities copied straight from a tool with no impact analysis. |
5. Who owns the engagement day to day, and how do we reach the testers, not just an account manager? | A named technical lead, a defined communication channel, and direct tester access during the test window. | All contact routed through sales, with no access to the people doing the work. |
Remaining general and methodology questions to include, scored on the same 0 to 5 scale:
What certifications and accreditations does your firm hold, and which are firm-level versus individual?
How long has the firm been delivering this exact type of testing, and can you share relevant references?
How do you handle scope changes discovered mid-test, and what is the change-control process?
What is your typical timeline from kickoff to draft report for an engagement of our size?
How do you avoid and disclose conflicts of interest, including reselling the products you assess?
What does your kickoff require from us, and how do you keep the engagement from stalling on our side?
How do you measure the quality of your own testing, and what stops a weak test shipping as a strong report?
Section 2: Tester qualifications and red team (weight 20)

This is the highest-weighted section because the people matter more than the process. A named, credentialed tester with years of relevant experience finds the bugs that decide the engagement. This section also carries the regulator-grade questions: if you are a financial entity in scope for the EU's Digital Operational Resilience Act, or you run intelligence-led testing under the UK's CBEST scheme, the tester bar is not a preference, it is written into the rules.
The DORA threat-led penetration testing Regulatory Technical Standards, applicable since 8 July 2025, require external testers to field a lead with at least five years of penetration and red team testing experience, plus at least two more testers with at least two years each, a combined five or more prior assignments, at least five references, curricula vitae with market-standard certifications, and professional indemnity insurance (EU 2025/1190, Article 7). The ECB's TIBER-EU procurement guidance sets a parallel bar: a red team manager with at least five years of experience including three years leading tests in financial services, and each red team member with at least two years (TIBER-EU, ECB). Any buyer can lift these thresholds into an RFP as pass-or-fail criteria, even outside financial services.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. Who specifically will test our environment, and what are their qualifications? | Named testers with relevant certifications (OSCP, OSCE3, OSWE, CREST CRT, CRTO) and years of hands-on experience on similar targets. | "We will assign qualified staff" with no names, no certifications, and no CVs. |
2. Are the testers your own staff or subcontractors, and where are they located? | In-house testers, or disclosed and vetted subcontractors, with locations relevant to your data-residency needs. | Undisclosed subcontracting, or refusal to say who actually performs the work. |
3. For a red team engagement, what is your adversary-emulation methodology and how do you map to MITRE ATT&CK? | Threat-intelligence-led scenarios, explicit ATT&CK technique mapping, and objectives tied to your crown-jewel assets. | A vulnerability scan relabeled as a red team, with no adversary model or ATT&CK mapping. |
4. For regulated testing (DORA TLPT, TIBER-EU, CBEST), how do your testers meet the scheme's qualification requirements? | Specific mapping to the scheme: lead experience years, references, certifications, insurance, and separation of threat intelligence from red team roles. | Vague "we are compliant" with no reference to the actual tester-qualification thresholds. |
5. How do you separate the threat-intelligence function from the red team, and why does that matter? | Independent intelligence provider or a walled-off internal team, so scenarios reflect real adversaries rather than the red team's convenience. | Threat intelligence and red team collapsed into one role with no independence. |
6. What are the rules of engagement, deconfliction, and stop conditions for a live red team test? | A signed rules-of-engagement document, named deconfliction contacts, immediate stop authority, and legal authorization to test. | No written rules of engagement, or no clear way to halt the operation. |
Remaining tester and red team questions to include:
What is the average tenure and seniority of the testers you will assign to us?
How do you keep testers current: continuing certification, internal research time, CVE publication, conference work?
Can we interview or veto the proposed lead tester before the engagement starts?
For physical or social-engineering components, what authorization and safety controls do you require in writing?
How do you handle a tester leaving mid-engagement without losing continuity or context?
What is your process for responsible disclosure if you find a zero-day in third-party software during our test?
How do you demonstrate independence from the vendors and products in our environment?
What insurance and liability cover do you carry, and can you provide a certificate?
For grading the red team portion in depth, pair this section with grading MITRE ATT&CK coverage in a red team proposal and the red team rules-of-engagement buyer checklist.
Section 3: Cloud and infrastructure depth (weight 15)
Cloud RFP criteria are where generic pentest shops thin out fastest. Testing an AWS, Azure, or GCP environment properly means reviewing identity and access management, control-plane configuration, and workload isolation, not just running a network scan against public IPs. This section is worth 15 because a shallow cloud test gives false comfort in exactly the layer where most modern breaches now start.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. How do you test cloud IAM: roles, trust policies, privilege escalation paths, and cross-account access? | Manual IAM policy review plus escalation-path testing (for example, role chaining and confused-deputy issues), with named cloud experience. | "We scan the external IP range" with no IAM or control-plane testing at all. |
2. How do you assess the cloud control plane and configuration, not just the workloads? | Review of the management plane against a named benchmark (CIS, provider well-architected guidance), with misconfiguration testing. | Configuration review skipped, or reduced to a generic checklist with no manual validation. |
3. How do you test container and Kubernetes security: escapes, RBAC, secrets, and workload isolation? | Concrete cluster test cases: pod escape attempts, RBAC review, secrets handling, and namespace isolation checks. | Kubernetes treated as out of scope or "the same as any Linux host." |
4. What access model do you use for cloud testing, and how do you protect the credentials we provide? | Least-privilege scoped roles, short-lived credentials, and a clear handling and destruction process for anything we issue. | Requests for long-lived admin keys with no explanation of how they are protected or revoked. |
Remaining cloud and infrastructure questions to include:
How do you test serverless functions, managed services, and provider-specific attack surface?
How do you approach hybrid and on-premise-to-cloud trust boundaries?
Do you test the CI/CD pipeline and infrastructure-as-code, and how?
How do you avoid triggering cloud-provider abuse detection or disrupting shared tenancy?
What cloud-specific certifications or hands-on experience do your assigned testers hold?
How do you validate network segmentation and lateral movement inside a virtual private cloud?
How do you handle testing across multiple accounts, subscriptions, or projects in one engagement?
Section 4: AI and automation disclosure (weight 15)

This section did not exist in most RFP templates a year ago, and it is now non-negotiable. In Cobalt's AI and Pentesting Pulse Report 2026, a survey of 455 security leaders and practitioners, support for fully automated pentesting fell from 29 percent to 9 percent year over year, 78 percent of organizations reported fully automated scanning tools missing critical vulnerabilities and returning false negatives, and 47 percent now prefer a hybrid model where humans support AI testing (Cobalt, 2026). Adoption of AI in testing is still climbing, HackerOne recorded a 210 percent year-over-year jump in valid AI-assisted vulnerability reports (HackerOne, 2025), so the question is not whether a vendor uses AI, it is whether they disclose it and validate it.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. Do you use AI or automated agents in testing, and which findings come from them versus a human? | Transparent disclosure of where AI is used, with human validation of every high-severity finding before it reaches your report. | "No" that later turns out to be untrue, or "everything is human" with no way to verify. |
2. How do you prevent AI-driven false positives from reaching our report? | A named validation step: a senior tester confirms each finding, and false-positive rate is measured, not asserted. | Raw AI output passed through with no human triage, which is exactly the source of the 78 percent false-negative problem above. |
3. If you run an autonomous agent against our systems, what runtime controls and scope enforcement are in place? | Enforced scope, rate limits, an immediate stop control, and a full audit trail of every action the agent took. | An autonomous tool pointed at production with no scope enforcement or kill switch. |
4. How is our data used with any AI tooling: is it sent to third-party models, and is it used for training? | Clear data-flow disclosure, no training on your data, and named model or self-hosted arrangement with contractual backing. | Vague "we use AI" with no answer on where your data goes or whether it trains a third-party model. |
5. What can your AI find that a scanner cannot, and where does it still need a human? | Honest scope: strong on coverage and speed, with humans required for business logic, chained exploits, and impact validation. | A claim that AI fully replaces human testers, which no current benchmark supports. |
Remaining AI and automation questions to include:
What is the measured false-positive and false-negative rate of your automated tooling, and how do you measure it?
How do you keep AI tooling from acting outside the agreed scope during an unattended run?
Do you offer a hybrid model, and how is human review structured within it?
How do you handle prompt-injection and data-exfiltration risk in your own AI tooling?
Can you run a small proof of concept on our own application before we commit?
How do you version and log AI tool behavior so a finding is reproducible months later?
What happens to AI-generated evidence and logs at the end of the engagement?
For a deeper vendor-side interrogation of AI claims, use the questions to ask an AI pentest vendor, and to prove the claims on your own app, run the 30-day AI pentest bake-off scorecard.
Section 5: Reporting quality (weight 15)
The report is the deliverable you actually keep. A strong report is reproducible, prioritized by real risk, and usable by both an engineer fixing the bug and an auditor reviewing the evidence. Ask for a sample before you sign, because a redacted real report tells you more than any proposal paragraph.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. Can you provide a redacted sample report from a comparable engagement? | A real, redacted report with clear reproduction steps, evidence, severity justification, and remediation guidance. | Refusal to share any sample, or a one-page "certificate" with no technical detail. |
2. What does a single finding contain, from reproduction steps to remediation? | Reproduction steps, proof-of-concept evidence, business impact, CVSS or rated severity, and specific, testable remediation. | Findings that are just scanner output with a generic "update to latest version" fix. |
3. How do you tie findings to standards and compliance evidence we can hand to an auditor? | Mapping to the frameworks you name (SOC 2, ISO 27001, PCI DSS, DORA), so the report doubles as audit evidence. | No mapping, leaving you to translate raw findings into audit language yourself. |
4. What do you deliver beyond the PDF: readout, developer walkthrough, machine-readable output? | A findings review call, a developer-facing walkthrough, and optional structured output for your ticketing system. | Report emailed with no debrief and no path to the people who have to fix the issues. |
Remaining reporting questions to include:
How do you prioritize findings when everything cannot be fixed at once?
How do you handle false positives found after delivery, and do you correct the report?
What is your turnaround from end of testing to draft report, and to final report?
Do you provide an executive summary that a non-technical stakeholder can act on?
How do you present findings that are chained, where two medium issues combine into a critical?
Will you present to our board or auditors if we need you to, and at what cost?
For a full rubric on grading the deliverable itself, see how to evaluate a penetration test report. If the report has to carry compliance weight, SOC 2 penetration testing and PCI DSS penetration testing set out what each framework expects the test itself to cover.
Section 6: Retesting and remediation (weight 10)
A finding you cannot prove is fixed is a finding you cannot close. Retesting is where thin bids quietly save money, by excluding fix verification, so the true cost lands on you the first time an auditor asks for evidence of remediation. Ask exactly what retesting is included and what triggers a paid re-scope.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. Is retesting of fixed findings included, and for how long after delivery? | Retesting included within a defined window (commonly 30 to 90 days), with a clear statement of what counts as a fix verification versus a new test. | Retesting billed as a full new engagement, or no retest offered at all. |
2. How do you verify a fix: targeted re-test of the finding, or a full re-scan? | A targeted re-test that confirms the specific issue is resolved without re-charging for the whole engagement. | Only a full re-scan, which inflates cost and slows down your remediation cycle. |
3. How do you support our developers during remediation, and is that time included? | Access to the tester for remediation questions, and clear guidance on fixes, within the engagement scope. | Remediation support treated entirely as billable extra with no included support at all. |
Remaining retesting and remediation questions to include:
What is your SLA for confirming a critical-severity fix?
How do you track remediation status across multiple findings over time?
Do you offer continuous or periodic retesting, and how is that priced against a one-off test?
How do you handle a finding we dispute or accept as a risk rather than fix?
What happens if a fix introduces a new vulnerability that your retest uncovers?
For the contractual side of retest guarantees and turnaround, see autonomous pentest contract and SLA clauses.
Section 7: Data handling and legal (weight 10)
This section protects you from turning a security test into a data-protection incident. You are granting a third party access to your systems and, often, your data, so the RFP has to pin down where that data lives, how long it is kept, and who is liable if the engagement goes wrong.
Question | A strong answer contains | Red-flag answer |
|---|---|---|
1. Where is our data stored during and after the engagement, and for how long is it retained? | Named storage locations, a defined retention period, and a documented secure-destruction step with confirmation. | No clear answer on where data lives, or indefinite retention with no destruction policy. |
2. What contractual protections do you offer: NDA, liability cover, data-processing terms? | A signed NDA, professional indemnity insurance, and data-processing terms that match your regulatory obligations. | Reluctance to sign an NDA, or no liability cover and no data-processing agreement. |
3. How do you ensure testing does not disrupt production or expose real customer data? | Non-destructive testing rules, staging where appropriate, and explicit handling for any real data encountered. | No safeguards for production, or a willingness to test destructively without written authorization. |
Remaining data handling and legal questions to include:
How do you securely transmit findings and evidence to us?
Who on your side has access to our data, and how is that access controlled and logged?
How do you handle a breach of your own systems that could expose our engagement data?
What is your subprocessor list, and how are subprocessors bound to the same terms?
How do you support our own regulatory obligations (GDPR, sector rules) as a data processor?
The regulated-sector overlay: what changes when a bank runs the RFP
In regulated financial services, several of the criteria above stop being preferences and become supervisory requirements. Treat them as pass-or-fail prequalification rather than scored questions, because a threshold set by a regulator is not something a bidder can win back on price. Add the regimes that apply to you, keep the seven scored sections exactly as they are, and your total still works out of 100.
Regime | Applies to | What it fixes in your RFP | Primary source |
|---|---|---|---|
DORA TLPT | EU financial entities in scope for threat-led penetration testing | Tester and threat-intelligence qualification thresholds, references, certifications, insurance | Commission Delegated Regulation (EU) 2025/1190, Article 7 |
TIBER-EU | EU financial entities running intelligence-led red team tests | Red team manager and member experience, separation of threat intelligence from red team | TIBER-EU framework and procurement guidance, ECB |
CBEST | UK financial firms in scope for intelligence-led assessment | Accreditation and certification expectations for providers | CBEST implementation guide, Bank of England |
APRA CPS 234 | APRA-regulated banks, insurers, and superannuation entities in Australia | Systematic testing program, tester independence, third-party testing, board escalation | CPS 234, paragraphs 27 to 36 |
MAS TRM Guidelines | MAS-regulated financial institutions in Singapore | Testing scope and frequency, adversarial attack simulation, intelligence-based scenario design | MAS Technology Risk Management Guidelines, Chapter 13 |
DORA TLPT and TIBER-EU: the tester bar is written into the rules (EU)
Section 2 covers this in detail. The short version for prequalification: Article 7 requires external testers to field a lead with at least five years of penetration and red team testing experience plus at least two more testers with at least two years each, combined participation in at least five previous penetration testing assignments, at least five references, CVs with market-standard certifications, and full professional indemnity insurance. The same article sets a parallel bar for the threat-intelligence provider: a manager with at least five years of threat-intelligence experience, at least one further member with at least two years, at least three references, and combined participation in at least three previous threat-intelligence assignments (EU 2025/1190, Article 7). Write both bars into the prequalification, because a bidder can clear the red team bar and fail the intelligence bar.
APRA CPS 234: skilled and functionally independent testers (Australia)
CPS 234 gives an Australian buyer the strongest single requirement in any of these regimes: testing must be conducted by "appropriately skilled and functionally independent specialists" (paragraph 30). Independence is the operative word, and it rules out the arrangement where the firm that built or operates a control is also the firm that tests it.
The rest of the standard translates directly into RFP clauses:
Paragraph 27 requires a systematic testing program whose nature and frequency is commensurate with the rate at which vulnerabilities and threats change, the criticality and sensitivity of the asset, the consequences of an incident, exposure to environments where you cannot enforce your own policies, and the materiality and frequency of change to the asset. Require each bidder to explain, in their own words, how the proposed scope and cadence address all five factors.
Paragraph 28 applies where a third party manages your information assets and you rely on that party's testing: you must assess whether the nature and frequency of that testing is commensurate with the same five factors. Require bidders to state how their evidence supports that assessment.
Paragraph 29 requires escalation to the board or senior management of any testing result identifying control deficiencies that cannot be remediated in a timely manner. Require reporting that flags exactly those findings rather than burying them inside a severity table.
Paragraph 31 requires review of the sufficiency of the testing program at least annually, or when there is a material change. That clause is what justifies a term contract with a scheduled cadence instead of an ad hoc purchase.
Paragraphs 35 and 36 require notification to APRA no later than 72 hours after becoming aware of a material information security incident, and no later than 10 business days after becoming aware of a material control weakness the entity does not expect to remediate in a timely manner. Require bidders to commit to notification timelines that let you meet yours.
For the Australian provider shortlist, see top penetration testing companies in Australia.
MAS TRM Guidelines: black box plus grey box, annually, on production (Singapore)
Chapter 13 of the MAS Technology Risk Management Guidelines is unusually specific, which makes it easy to lift into a document:
13.2.1 expects a combination of black box and grey box testing for online financial services. Write both into the scope matrix, and require them priced as separate lines rather than assumed into one.
13.2.3 expects testing on the production environment with proper safeguards implemented. Require the bidder's written production safeguards with the bid, not as a kickoff conversation.
13.2.4 expects penetration testing of internet-facing systems at least once annually, or whenever those systems undergo major changes or updates. That second clause is exactly why a change trigger belongs in a continuous testing contract.
13.4.1 and 13.4.2 expect an adversarial attack simulation exercise, with objectives, scope, and rules of engagement defined before commencement, run under close supervision so red team activity does not disrupt production systems.
13.5 expects intelligence-based scenario design: scenarios built on challenging but plausible threats, using threat intelligence relevant to your own environment to identify the threat actors most likely to target you and the tactics, techniques, and procedures they use.
13.6 expects a remediation process that tracks and resolves the issues found, including severity assessment and classification. Require the bidder's retest and tracking model to plug into it.
For the APAC provider shortlist, see top penetration testing companies in Singapore.
Regulated-sector prequalification checklist
Run this before scoring anything. A bidder that fails any applicable line does not reach the scored sections.
The bidder meets the tester-qualification thresholds of the regime that applies to us, evidenced with CVs and certification numbers.
The threat-intelligence function is independent of the red team, and the bidder can show us how that separation works.
The bidder is functionally independent of the controls being tested, with no role in building or operating them.
Firm-level accreditation is current and verified in the accrediting body's directory, for the discipline we are actually buying.
Professional indemnity insurance is in force at the level our regime expects, evidenced by certificate.
Rules of engagement, deconfliction contacts, and stop authority are agreed in writing before any activity begins.
Reporting supports our board escalation and our supervisory notification timelines.
Engagement data residency and retention match our regulatory obligations.
A worked set of strong answers
The fastest way to calibrate the red-flag key is to read a set of strong answers end to end. As the buyer, you are entitled to hold your shortlist to this standard. Here is how a hybrid, human-plus-AI provider answers a handful of the hardest questions in this bank, written the way you should expect a serious bidder to answer.
On firm accreditation and tester qualifications (Section 2, Q1 and Q4). Stingrai is a CREST-accredited firm, and its engagements are delivered by senior human pentesters whose certifications include OSCE3, OSCP, OSWE, CREST CRT, and CRTO, alongside 18 published CVEs and research presented at DEF CON and BSides. Testers are named before the engagement, and for regulated testing the qualification mapping is explicit rather than a blanket compliance claim.
On methodology and manual depth (Section 1, Q1 to Q3). The work runs on a named methodology, and reporting is mapped to MITRE ATT&CK so red team coverage is legible rather than asserted. Manual effort concentrates on the classes automated tools miss: broken authorization, IDOR, and business-logic flaws.
On AI disclosure and the hybrid model (Section 4, Q1 and Q5). Stingrai delivers a transparent hybrid model. Snipe, its autonomous web-application testing agent, is purpose-built to hunt the complex classes generic scanners miss, IDOR, broken access control, and business-logic flaws, and it works black-box and white-box, reviewing source code, opening AutoFix pull requests, and running as a pull-request gate that can block vulnerable code from merging. It is trained on more than 6,000 HackerOne Hacktivity disclosure reports and on skills distilled from the firm's own human pentesters. Cloud, identity, network, and social-engineering testing is led by human pentesters. Every high-impact finding is validated by a senior human before it reaches your report, which is the step that keeps false positives off the page.
On retesting and evidence (Section 5 and Section 6). Retesting and AutoFix pull requests are part of the delivery model rather than a billable afterthought, and the reporting is built to support your SOC 2, ISO 27001, PCI DSS, and DORA evidence needs so the deliverable does double duty as audit material. Pricing for these packages is published openly at stingrai.io/pricing.
You do not have to shortlist any particular vendor. You do have to hold whoever you shortlist to answers at this level of specificity. If a bidder cannot, the red-flag key will tell you in a single scoring pass. To see the full service scope behind these answers, review web application penetration testing, red teaming, and the wider penetration testing services.
Frequently Asked Questions
What questions should I include in a penetration testing, red team, or cloud pentest RFP, and what vendor answers are red flags?
Include questions across seven areas: general methodology, tester qualifications and red team capability, cloud and infrastructure depth, AI and automation disclosure, reporting quality, retesting, and data handling. The clearest red flags are a scanner-only workflow sold as a manual pentest, testers the vendor will not name or credential, no sample report, findings with no retest included, an autonomous AI tool run against production with no scope enforcement, and any refusal to say which findings a human validated. This bank scores all seven areas on a 0 to 5 scale so each red flag becomes a number.
What criteria should be included in an RFP for an enterprise autonomous pentest?
An enterprise autonomous pentest RFP should score nine criteria: runtime scope enforcement with a working stop control, human validation of every high-severity finding before it reaches your report, measured false-positive and false-negative rates, evidence of which vulnerability classes the agent reaches beyond the known classes a scanner already finds, disclosure of where your data goes and whether it trains a third-party model, a reproducible audit trail of every action the agent took, integration with your ticketing and pull-request workflow, a per-asset or subscription commercial model with a defined exit, and a proof of concept against one of your own applications before you commit. Section 4 of this bank turns each criterion into a scored question with its red-flag answer.
What should a bank put into an RFP for red team services?
A bank should build its red team RFP around four things a general pentest RFP leaves out: the supervisor's tester-qualification thresholds written in as pass-or-fail prequalification, enforced independence between the threat-intelligence function and the red team, a rules-of-engagement and deconfliction regime that protects production and customer data, and objectives defined against named critical business functions rather than a list of IP ranges. Add measurement of your own defenders, because the deliverable that matters in a regulated firm is evidence about detection and response rather than a vulnerability list. The thresholds themselves come from DORA TLPT and TIBER-EU in the EU, CBEST in the UK, APRA CPS 234 in Australia, and the MAS adversarial attack simulation expectations in Singapore.
How do you put together an RFP for continuous penetration testing?
Define what continuous means before you ask anyone for a price, because the word currently covers three different products: always-on automated scanning, scheduled waves of human testing spread across the year, and testing gated on every code change. State which one you are buying and the human testing hours included per period, then specify the asset inventory and the change triggers that pull an asset back into testing, the cadence and how the vendor proves coverage, the retest turnaround included in the subscription, delivery of findings into your ticketing system, a periodic summary report your auditors can use as evidence that testing ran across the term, and commercial terms covering estate growth, notice period, and exit with confirmed data destruction.
What sections should a penetration testing RFP template contain?
A complete penetration testing RFP template contains nine parts: company background and firm qualification, a scope definition matrix, methodology and standards requirements, tester qualification requirements with a verification procedure, evidence and reporting requirements, service levels with retesting and contract terms, a commercial response format, the evaluation and scoring rubric you publish to bidders, and the automatic disqualifiers that end a bid. Add submission instructions and a timeline, then paste your scored technical questionnaire into the methodology and tester qualification parts so bidders answer it inside the document.
How do you verify a penetration testing company's CREST accreditation?
Verify it in five steps. Ask for the exact legal entity name that holds the accreditation and check it matches the entity that will sign your contract, because group structures often hold accreditation in one subsidiary and deliver the work from another. Search the accrediting body's public member directory rather than trusting a logo on a website, since CREST publishes its member companies at marketplace.crest.org. Confirm which discipline the accreditation covers, because it is granted per service rather than to the whole firm. Separate firm-level accreditation from individual certifications and require both, evidenced separately. Finally, ask for the current status and renewal date, and require written notice if it lapses during the term.
What does APRA CPS 234 require in a penetration testing RFP?
CPS 234 requires that testing be conducted by appropriately skilled and functionally independent specialists (paragraph 30), which rules out a firm testing controls it built or operates. Paragraph 27 requires a systematic testing program whose nature and frequency is commensurate with the rate at which threats change, the criticality and sensitivity of the asset, the consequences of an incident, exposure to environments you cannot enforce policy over, and the rate of change to the asset. Paragraph 29 requires board escalation of deficiencies that cannot be remediated in a timely manner, paragraph 31 requires review of the testing program's sufficiency at least annually, and paragraphs 35 and 36 set APRA notification at no later than 72 hours for a material incident and 10 business days for a material control weakness.
What do the MAS TRM Guidelines require for penetration testing and red teaming?
Chapter 13 expects a combination of black box and grey box testing for online financial services (13.2.1), testing on the production environment with proper safeguards implemented (13.2.3), and penetration testing of internet-facing systems at least once annually or whenever those systems undergo major changes or updates (13.2.4). For red teaming, 13.4 expects an adversarial attack simulation exercise with objectives, scope, and rules of engagement defined before commencement and run under close supervision so production is not disrupted, 13.5 expects intelligence-based scenario design using threat intelligence relevant to your environment, and 13.6 expects a remediation process that tracks and classifies the issues found.
What is a penetration testing RFP template?
A penetration testing RFP template is a structured request for proposal that a buyer sends to security testing vendors to compare them on the same criteria. A strong template goes beyond a question list: it weights the sections, scores each answer, and defines the red-flag answers in advance, so replies from different vendors are genuinely comparable instead of a stack of confident prose. This bank is a scored template you can copy directly into your procurement process.
How do you score penetration testing vendor responses in an RFP?
Weight the sections to your own risk so they total 100, then score each answer from 0 to 5, where 0 is a documented red-flag answer, 3 is adequate, and 5 is the strong answer defined in the key. Multiply each section's average by its weight and total to 100. Add an override rule: a red-flag answer on any high-weight question caps that section regardless of the rest of the reply, which stops a strong sales narrative from hiding a fundamental gap.
What are red flags in a penetration testing vendor's proposal?
The most common red flags are automation sold as manual testing, unnamed or uncredentialed testers, no redacted sample report, business logic and authorization left out of scope, findings delivered with no retest, and vague answers about where your data goes and how long it is kept. In AI-assisted testing, the added red flags are raw automated output with no human validation and any autonomous agent pointed at production without enforced scope and a stop control.
What red team tester qualifications should a DORA TLPT or TIBER-EU RFP require?
Lift the thresholds straight from the rules. The DORA threat-led penetration testing Regulatory Technical Standards (EU 2025/1190, Article 7) require a lead tester with at least five years of penetration and red team testing experience, at least two more testers with at least two years each, a combined five or more prior assignments, at least five references, CVs with market-standard certifications, and professional indemnity insurance. The ECB's TIBER-EU guidance adds a red team manager with at least five years including three years leading financial-sector tests, and independence between the threat-intelligence and red team functions.
What AI and automation disclosure questions should a pentest RFP include?
Ask whether the vendor uses AI or autonomous agents, which findings come from automation versus a human, how false positives are prevented, what runtime scope controls and stop conditions apply to any autonomous agent, and whether your data is sent to third-party models or used for training. This matters because support for fully automated pentesting fell from 29 percent to 9 percent in a year and 78 percent of teams hit critical false negatives from automated tools (Cobalt, 2026). Disclosure and human validation, not the presence of AI, are what you are scoring.
What cloud pentest RFP criteria matter most?
The criteria that separate a real cloud test from a network scan are identity and access management testing (roles, trust policies, privilege escalation), control-plane and configuration review against a named benchmark, container and Kubernetes security, and a least-privilege access model for any credentials you issue. A vendor that answers "we scan the external IP range" for a cloud environment is testing the wrong layer, because most cloud compromise starts in identity and configuration, not the network edge.
What should a pentest RFP ask about reporting quality?
Ask for a redacted sample report before signing, and confirm that each finding contains reproduction steps, evidence, business impact, a rated severity, and specific remediation. Ask how findings map to the compliance frameworks you care about, and what you get beyond the PDF: a findings review call, a developer walkthrough, and machine-readable output. A one-page certificate with no technical detail is a red flag, because the report is the deliverable you keep and hand to auditors.
What should a pentest RFP ask about retesting?
Ask whether retesting of fixed findings is included and for how long, whether a fix is verified with a targeted re-test or a full re-scan, and what SLA applies to confirming a critical fix. Retesting is where thin bids hide cost by excluding fix verification, so pin it down in the RFP. A vendor that only offers a full new engagement to confirm a fix will slow your remediation and inflate your spend.
What questions should a pentest RFP include about data handling?
Ask where your data is stored during and after the engagement, how long it is retained, and how it is securely destroyed. Confirm the vendor will sign an NDA, carries professional indemnity insurance, and offers data-processing terms that match your regulatory obligations. Ask who on their side can access your data, how that access is logged, and what their subprocessor list looks like. Vague retention answers or reluctance to sign an NDA are red flags.
References
Cobalt. AI and Pentesting Pulse Report 2026. June 2026. https://resource.cobalt.io/ai-pentesting-pulse-report-2026-tyd. Survey of 455 security leaders and practitioners built on five years of pentesting data; source for the 29 percent to 9 percent drop in support for fully automated pentesting, the 78 percent critical-false-negative figure, and the 47 percent hybrid-model preference.
HackerOne. 2025 Hacker-Powered Security Report. November 2025. https://www.hackerone.com/blog/ai-security-trends-2025. Source for the 210 percent year-over-year jump in valid AI-assisted vulnerability reports.
European Commission. Commission Delegated Regulation (EU) 2025/1190 (DORA TLPT Regulatory Technical Standards). Applicable 8 July 2025. https://eur-lex.europa.eu/eli/reg_del/2025/1190/oj. Source for threat-led penetration testing tester-qualification requirements in Article 7: lead experience, additional testers, references, certifications, and insurance.
European Central Bank. TIBER-EU Framework and Service Provider Procurement Guidance. https://www.ecb.europa.eu/paym/cyber-resilience/tiber-eu/html/index.en.html. Source for red team manager and red team member experience thresholds and independence between threat intelligence and red team roles.
Bank of England. CBEST Threat Intelligence-Led Assessments Implementation Guide. https://www.bankofengland.co.uk/financial-stability/operational-resilience-of-the-financial-sector/cbest-threat-intelligence-led-assessments-implementation-guide. Source for CREST accreditation and certification expectations governing intelligence-led financial-sector testing.
CREST. Certifications and accreditation. https://www.crest-approved.org/. Source for the red team and threat-intelligence certifications referenced in the tester-qualification section.
Australian Prudential Regulation Authority. Prudential Standard CPS 234 Information Security. July 2019. https://www.apra.gov.au/sites/default/files/cps_234_july_2019_for_public_release.pdf. Source for the systematic testing program requirement (paragraph 27), third-party testing assessment (paragraph 28), board escalation of un-remediated deficiencies (paragraph 29), the requirement for appropriately skilled and functionally independent specialists (paragraph 30), annual review of testing sufficiency (paragraph 31), and the 72-hour and 10-business-day notification obligations (paragraphs 35 and 36).
Monetary Authority of Singapore. Technology Risk Management Guidelines. January 2021. https://www.mas.gov.sg/regulation/guidelines/technology-risk-management-guidelines. Source for the Chapter 13 cyber security assessment expectations: black box and grey box testing for online financial services (13.2.1), production testing safeguards (13.2.3), annual penetration testing of internet-facing systems (13.2.4), adversarial attack simulation exercises (13.4), intelligence-based scenario design (13.5), and remediation management (13.6).
CREST. CREST Marketplace member directory. https://marketplace.crest.org. Source for verifying that a bidding company holds a current firm-level accreditation, and for which discipline that accreditation covers.
Score your shortlist against real answers
You now have a copy-ready RFP template, a scored question bank, a red-flag answer key, and the regulator-grade thresholds to hold vendors to. The last step is comparing the replies against a provider that answers at this level of specificity. Stingrai, a CREST-accredited firm, delivers hybrid penetration testing and red teaming that pairs senior human testers with Snipe, its autonomous web-application agent, with named methodology, MITRE ATT&CK-mapped reporting, AutoFix pull requests, included retesting, and evidence built to support your SOC 2, ISO 27001, PCI DSS, and DORA programs. See the penetration testing services, the PTaaS platform, and open pricing to benchmark your shortlist.
Take the template, issue it to whoever you like, and hold every bidder to the same rubric. When you are ready for responses, send your RFP to Stingrai or request a quote and we will answer it question by question, in the format above.



