The Penetration Testing Execution Standard divides an engagement into seven sections, from pre-engagement interactions to reporting. NIST SP 800-115, the Technical Guide to Information Security Testing and Assessment, publishes a four-phase methodology instead: planning, discovery, attack and reporting, with a feedback loop from attack back into discovery. The OWASP Web Security Testing Guide is not a phase model at all but a catalogue of numbered web test cases. OSSTMM 3 is a measurement framework across five channels. All four are routinely called "the penetration testing methodology", and they are answering different questions. That confusion is why so many buyers cannot tell a rigorous test from a scan with a report template attached.
Quick answer: A penetration testing methodology is the documented, repeatable process a tester follows so that the coverage of a test is provable rather than asserted, and in practice a credible one names several standards at once: PTES or NIST SP 800-115 for phase structure, OWASP WSTG and ASVS for web coverage, OWASP MASVS and MASTG for mobile, OWASP API Security Top 10 for APIs, and MITRE ATT&CK where the work is adversary simulation. Stingrai is a CREST-accredited offensive security company, and its penetration testers, from a team holding OSCE3, OSCP, OSWE, CREST CRT and CISSP, run engagements against exactly that stack across applications, cloud, networks and people. Published packages cover one web application and its APIs at US$3,000 one-time or US$650 per month on a 12-month continuous engagement, listed on the pricing page; network, cloud, Active Directory, wireless, physical and social engineering engagements are quoted by scope through get a quote.

What a penetration testing methodology actually is
A methodology is not a tool list and it is not a marketing page. It is the answer to three questions a reader of your report will ask a year later, when the person who commissioned the test has left and an auditor or an acquirer is reading it cold:
What was in scope, and how was that decided?
What was actually attempted against it, and in what order?
What evidence exists that each step happened?
Everything else follows from those. A test with a documented methodology can be repeated next year and compared. A test without one produces a list of findings that may be complete, may be a scanner export, and cannot be distinguished from the outside.
The PCI Security Standards Council makes this explicit in its Penetration Testing Guidance information supplement, version 1.1, September 2017. Its report-evaluation checklist asks two questions before any finding is read: "Is the methodology clearly stated?" and "Does the methodology reflect industry best practices (OWASP, NIST, etc.)?" The same supplement requires that a report contain a Statement of Methodology and a Statement of Limitations as named sections. That is the standard a buyer should hold every vendor to, whether or not card data is involved.
Two things a methodology is not. It is not a guarantee of findings, because a well-hardened target can pass. And it is not a substitute for the tester, because every standard on this page leaves the hard judgement calls, which attack path to pursue and when a chain is worth exploiting, to the person running the test.
The standards map: what each one governs

Standard | Published by | What it governs | Where it stops |
|---|---|---|---|
PTES | Community project, v1.0 | End-to-end execution, in seven sections | No technical detail in the standard itself; that sits in a separate technical guide |
NIST SP 800-115 | NIST, September 2008 | Four-phase technical testing methodology plus a rules-of-engagement template | Predates cloud, mobile and modern identity attacks |
OWASP WSTG | OWASP, v4.2 (3 December 2020) | Web application and web service test cases, individually numbered | Web only, and it does not set a pass or fail bar |
OWASP ASVS | OWASP, v5.0.0 (30 May 2025) | Pass or fail verification requirements across three levels | Says what to satisfy, not how to test it |
OWASP MASVS and MASTG | OWASP Mobile Application Security | Mobile control groups and the tests that verify them | Mobile clients and their platform interactions |
OWASP API Security Top 10 | OWASP, 2023 edition | The API risk classes a test must reach | A risk list, not a full test catalogue |
OSSTMM 3 | ISECOM, 2010 | Measurement: five channels, six test types, the rav attack-surface metric | Steep learning curve, and rarely used verbatim commercially |
MITRE ATT&CK | MITRE, v19.2 (28 April 2026) | Adversary tactics and techniques observed in the wild | A behaviour catalogue, not an engagement process |
CREST | CREST, not-for-profit | Accreditation of the supplier and certification of individuals | Assures who tests, not what a given test covered |
NCSC CHECK | UK NCSC | Assured testing for UK public sector and critical national infrastructure | Scheme eligibility is narrow by design |
PTES: the seven-section execution standard
PTES is the closest thing the industry has to a shared vocabulary for what an engagement contains. Its own front page states that the standard "consists of seven (7) main sections" covering everything from the initial communication and reasoning behind a test through to reporting. In order, they are:
Pre-engagement Interactions
Intelligence Gathering
Threat Modeling
Vulnerability Analysis
Exploitation
Post Exploitation
Reporting
Two caveats a buyer should know. First, PTES deliberately carries no technical instructions; those live in a separate PTES Technical Guidelines document. Second, the wiki describes itself as v1.0 with a v2.0 "in the works", and its main page was last edited in August 2014. It remains the best available map of engagement structure, and it is not a living catalogue of technique. Any vendor citing PTES alone for a 2026 cloud or identity engagement is citing a scaffold, not a syllabus.
NIST SP 800-115: the four-phase technical guide
NIST SP 800-115 is a US federal publication from September 2008, and Section 5.2.1 sets out what Figure 5-1 labels the "Four-Stage Penetration Testing Methodology": planning, discovery, attack, reporting. The detail worth knowing is in the arrows rather than the boxes. Successful exploitation loops back into additional discovery, because access to a new host restarts information gathering from a new vantage point. And reporting is not the last phase in practice: SP 800-115 notes that reporting "occurs simultaneously with the other three phases", starting with the assessment plan or rules of engagement written during planning.
Two parts of SP 800-115 are still doing real work in 2026. Appendix B is a rules-of-engagement template, and a great many commercial ROE documents are recognisably descended from it. And the document draws a hard line between a scanner checking for the possible existence of a vulnerability and the attack phase verifying it by attempting exploitation, which remains the cleanest available articulation of why a penetration test is not a vulnerability assessment. If your framework question is really about that distinction, our breakdown of what compliance frameworks actually require goes further.
Its age is the caveat. SP 800-115 predates public cloud as a default deployment target, mobile as a primary channel, and the identity-centric attacks that dominate current intrusions. Use it for phase structure and rules of engagement, not for what to test.
OWASP WSTG: the web test catalogue
The Web Security Testing Guide is the reference a web tester actually opens. It is a catalogue of individual test cases with stable identifiers such as WSTG-INFO-02, grouped by category, and OWASP describes it as "a comprehensive guide to testing the security of web applications and web services". The stable release is v4.2, published 3 December 2020, with v5.0 in development as of September 2026.
Its value in a report is traceability. A findings table that carries WSTG identifiers next to each test performed lets a reader confirm coverage without trusting a summary paragraph. When a proposal says "OWASP methodology" and cannot produce identifiers, it usually means the OWASP Top 10 awareness document, which is a list of risk categories and was never a testing methodology.
OWASP ASVS: the verification target
ASVS answers a different question: not how to test, but what the application must satisfy. Version 5.0.0 was released on 30 May 2025 at Global AppSec EU Barcelona. Every ASVS requirement is written to produce a pass or fail decision, and requirements are allocated to three levels by a priority-based evaluation that weighs risk reduction against implementation effort.
Level 1 is the minimum set and holds roughly 20% of requirements. OWASP notes it is "not necessarily penetration testable by an external tester without internal access to documentation or code", which is a useful corrective to the assumption that a black box test can certify an ASVS level.
Level 2 is where OWASP says most applications should be aiming. It carries around 50% of requirements, so reaching L2 means implementing roughly 70% of the standard.
Level 3 adds the final ~30%, largely defence-in-depth and harder controls, and is aimed at applications that need to demonstrate the highest assurance.
The practical pairing is ASVS as the target and WSTG as the method: ASVS says the application must do X, WSTG says how a tester confirms it. Buyers who want a verification statement rather than a findings list should name an ASVS level in the statement of work, and should expect to provide documentation or code access to make it reachable.
OWASP MASVS and MASTG: the mobile pair
The OWASP Mobile Application Security project mirrors that split. MASVS defines the control groups an application must satisfy, organised as MASVS-STORAGE, MASVS-CRYPTO, MASVS-AUTH, MASVS-NETWORK, MASVS-PLATFORM, MASVS-CODE, MASVS-RESILIENCE and MASVS-PRIVACY. MASTG is the testing companion, carrying platform-specific guidance for Android and iOS, test cases aligned to MASVS requirements, and demonstrations of real-world weaknesses.
The scoping consequence is that a mobile test is two tests. The client work is MASTG territory: storage, cryptography, platform interaction, resilience against reverse engineering. The backend the client talks to is a web and API engagement in its own right, and it is the half that usually produces the higher-severity findings. A mobile quote that does not name the backend is quoting half the work.
OWASP API Security Top 10: the risk classes an API test must reach
The 2023 edition is led by broken object level authorization and broken function level authorization, and both matter because neither is reliably reachable by an unauthenticated scanner. Testing them requires at least two accounts at different privilege levels and an understanding of which object belongs to whom, which is a scoping input, not a tool setting. We covered why tooling underperforms here in why API scanners miss BOLA and IDOR.
OSSTMM 3: measurement and channels
OSSTMM is the outlier, and the most misunderstood entry on this page. Published by ISECOM, version 3 dates from 2010 and describes itself as a methodology for "penetration and security testing, security analysis and the measurement of operational security". Three ideas from it have outlived the document's commercial adoption:
Five channels in three classes. OSSTMM divides interaction with assets into Human, Physical, Wireless, Telecommunications and Data Networks, grouped as PHYSSEC (physical and human), SPECSEC (wireless) and COMSEC (telecommunications and data networks). A full audit tests all five, and "we tested your network" quietly means one channel.
Six test types. Blind, Double Blind, Gray Box, Double Gray Box, Tandem and Reversal, distinguished by what the tester knows, what the target knows, and what the target expects. This is a far more precise vocabulary than the usual three boxes, because it separates the tester's knowledge from the defender's awareness.
The rav. A metric that scores the attack surface left open rather than counting findings, reported on a STAR (Security Test Audit Report) sheet.
Almost nobody delivers a full OSSTMM audit commercially. Its concepts, particularly the channel model and the tester-knowledge versus defender-awareness split, are worth borrowing into any scope document.
MITRE ATT&CK: the adversary behaviour catalogue
ATT&CK is "a globally-accessible knowledge base of adversary tactics and techniques based on real-world observations". The current release is v19.2, published 28 April 2026. The Enterprise matrix organises 274 techniques under 15 tactics, with separate Mobile and ICS matrices.
ATT&CK is not an engagement process, and treating it as one produces the most common proposal failure in red team buying: a bid that lists dozens of technique IDs with no statement of which ones will actually be attempted against which system. Used properly it does two jobs. It plans the exercise, via adversary emulation plans that MITRE describes as letting "red teams to more actively model adversary behavior, as described by ATT&CK". And it scores the result, by mapping each attempted technique to whether your controls prevented, detected or missed it. If you are evaluating an ATT&CK-flavoured proposal, our guide to grading ATT&CK coverage in a red team proposal is the checklist.
CREST and NCSC CHECK: assurance of the supplier, not the test
These two are frequently listed alongside methodologies and they belong in a different column. They assure who is testing, not what a particular test covered.
CREST is a not-for-profit accreditation body registered in the UK that operates two distinct tiers: accreditation of member companies for a service such as penetration testing, and certification of individuals by examination. Those are separate credentials and should not be conflated. A firm can hold company accreditation while its people hold individual certifications, and a buyer should ask which is being claimed.
NCSC CHECK is narrower by design. The NCSC describes it as setting "standards for penetration testing that government departments, public sector bodies and the UK's critical national infrastructure (CNI) organisations can trust". If you are not in one of those groups, CHECK eligibility is not the right filter for your shortlist; other assurance routes apply. Our directory of CREST-accredited penetration testing companies sets out what each credential actually evidences.
A real engagement, phase by phase, and the evidence each phase produces

Phase names are cheap. The test of whether a phase actually happened is the artefact it leaves behind. The table below is the one to put in front of a vendor, because it converts a methodology claim into a deliverables list.
Phase | What happens | Evidence it should produce |
|---|---|---|
Scoping and rules of engagement | Targets, boundaries, windows, constraints and authorisation are agreed and signed | Signed scope and target list, testing windows, escalation and deconfliction contacts, written authorisation to test |
Reconnaissance | Passive and active discovery of the real attack surface | A dated inventory of discovered assets against the assumed scope, with the source of each discovery |
Threat modelling | Attack paths are chosen against the assets that matter | The paths selected and the paths set aside, tied to business impact |
Vulnerability analysis | Candidate weaknesses are identified and triaged | Candidate list with identification method, and which automated output was manually confirmed |
Exploitation | Candidates are proven by attempting them | Requests and responses, screenshots, timestamps, and exact reproduction steps per finding |
Post-exploitation | Reach and impact of proven access are established | What data was provably readable, how far lateral movement went, what your detection stack saw |
Reporting | Findings are risk-rated and written up for two audiences | Risk ratings with reasoning, business impact, statement of methodology, statement of limitations |
Retest | Fixes are verified against the original finding | Per-finding verdict of fixed, partially fixed or still open, dated, plus an attestation letter |
1. Scoping and rules of engagement
This is PTES pre-engagement interactions and the NIST SP 800-115 planning phase, and it is where most bad engagements are already lost. The output is not a conversation, it is a document: in-scope targets by hostname, IP range, application URL, cloud account or site address; explicit exclusions; testing windows; techniques permitted and prohibited; data handling rules; a deconfliction contact reachable during testing; and signed authorisation from someone entitled to give it.
Cloud adds a layer here, because the platform provider sets rules you do not control. AWS lets customers test a named list of services "without prior approval" while prohibiting DoS, request flooding and bucket takeovers outright, and requires prior approval for anything involving command and control. Microsoft dropped its pre-approval requirement for Azure in June 2017 but points to the Microsoft Cloud Unified Penetration Testing Rules of Engagement as the authoritative text, which prohibits DDoS testing under all circumstances. A methodology that does not reference the provider rules for the platform in scope has not been written for cloud. Our walkthrough of cloud rules of engagement across AWS, Azure and GCP covers the differences.
If you are writing this document from scratch, start from our statement of work template, and if you are running a competitive process, the penetration testing RFP template turns the same content into questions vendors have to answer in writing.
2. Reconnaissance
PTES calls this intelligence gathering; NIST folds it into the first half of discovery. Passive work uses sources that never touch your infrastructure: certificate transparency logs, DNS records, code repositories, job listings, breach corpora. Active work touches it: port and service identification, banner grabbing, directory enumeration, cloud metadata discovery.
The evidence that matters here is the delta. A competent reconnaissance phase reliably returns assets the client did not have on its list, and the inventory of what was found against what was expected is often more valuable to a security team than any single finding. Dating it matters too, because a discovered asset is a point-in-time claim.
3. Threat modelling
This is the phase most often skipped and never noticed, because its absence looks identical to a busy test. Threat modelling decides which attack paths to spend the budget on. It asks what an attacker would want from this system, who would plausibly want it, and which chain of weaknesses gets them there.
The evidence is a record of choices: the paths pursued, and the paths deliberately not pursued with the reason. That second half is what makes a report honest. A test that spent four days on an authorisation model and none on denial of service should say so, rather than leaving a reader to assume uniform coverage.
4. Vulnerability analysis
Automated tooling belongs here, and so does the discipline of not trusting it. NIST SP 800-115 describes this as comparing the services, applications and operating systems found against vulnerability databases and the tester's own knowledge, and it is explicit that manual processes catch new or obscure issues that automated scanners miss while running far slower.
The evidence requirement is narrow and it is the single best question to put to a vendor: for each candidate issue, which tool or technique surfaced it, and was it manually confirmed. A finding that never got confirmation is a scanner output wearing a report's clothes.
5. Exploitation
This is the phase that separates a penetration test from everything else. A candidate is promoted to a finding by being proven, and the proof is what goes in the report: the request sent, the response received, the screenshot, the timestamp, and steps precise enough for a developer to reproduce the issue on their own machine.
Two rules keep this phase safe. Exploitation runs inside the boundaries set in phase one, so destructive techniques, denial of service and anything that touches third-party infrastructure are excluded unless explicitly authorised. And chained findings are reported as chains: three medium-severity issues that combine into full account takeover are one critical finding plus three mediums, not three mediums.
6. Post-exploitation
Access is not impact. Post-exploitation establishes what the access was actually worth: which data was provably readable, which additional systems were reachable, whether privileges could be escalated, and whether persistence was achievable within the rules.
One output from this phase is routinely under-requested and is the most useful thing in a modern report. Ask what your detection stack saw. A timeline of tester actions against alerts raised converts a penetration test into a free detection assessment, and if nothing fired, that is a finding in its own right. Where detection is the actual objective, a purple team exercise is the better purchase.
7. Reporting
Every standard on this page ends here, and NIST SP 800-115 points out that reporting really began in phase one with the assessment plan. A report serves two audiences that want opposite things: an executive summary in business terms, and per-finding technical detail an engineer can act on without a follow-up call.
The PCI SSC supplement names the sections a report should carry, including a Statement of Methodology covering the methodologies used, a Statement of Limitations documenting restrictions imposed on testing such as designated testing hours or legacy system constraints, a testing narrative describing how testing progressed and any interference encountered, and segmentation test results where segmentation is relied on. Those four are a good baseline whether or not PCI DSS applies to you. A worked example sits in our penetration testing report sample, and how to evaluate a penetration test report is the reviewer's checklist.
8. Retest
Retest is not in the classic phase models and it is the phase auditors ask about most. It verifies fixes against the original finding, produces a per-finding verdict of fixed, partially fixed or still open, and dates it. The deliverable is usually an attestation letter, which is what a customer security questionnaire or an auditor will actually want to see. Confirm at scoping whether retest is included and for how long the window stays open, because a retest bought later, as a new engagement, costs considerably more.
Black box, grey box and white box

Knowledge level is a scoping decision, not a quality ranking, and the most expensive mistake in this area is buying black box testing because it sounds more rigorous.
Black box. The tester begins with no more than a target name. It is the closest simulation of an unaided external attacker and it is the honest way to measure what your perimeter gives away. The cost is arithmetic: days spent on discovery are days not spent on depth, so coverage is the first casualty on a fixed budget.
Grey box. The tester gets partial knowledge: credentials for each user role, architecture notes, a documented API surface, a test account. This is the default for most application and network work and for almost all compliance-driven testing, because it puts the budget into the authorisation and business logic layers where the severe findings live. Its limitation is that it assumes a foothold rather than proving one can be won.
White box. Full knowledge, including source code, configuration and administrative credentials. It reaches depth nothing else does, particularly in authorisation logic, cryptographic implementation and anything requiring reasoning about intent rather than behaviour. It tells you the least about detection, and it demands the most preparation from your side.
OSSTMM 3 splits the same idea more precisely into six types, because it separates what the tester knows from what the target knows and expects. Blind and double blind describe tests the defender is not warned about; tandem and reversal describe fully cooperative arrangements. If you care whether your SOC is being tested as well as your systems, that vocabulary is worth adopting in your scope document.
How methodology changes with the target
The phase model holds across every engagement type. What changes is the reference catalogue, the scoping inputs and the evidence.
Target | Primary references | Scoping inputs that change the test | Evidence to demand |
|---|---|---|---|
Web application | OWASP WSTG, ASVS | Roles and credentials per role, authenticated vs unauthenticated surface, multi-tenancy | Test-case identifiers, per-role authorisation matrix results |
API | OWASP API Security Top 10, ASVS | Specification file, two accounts at different privilege levels, object ownership model | BOLA and BFLA attempts per endpoint, not per collection |
Mobile | OWASP MASVS, MASTG | Platforms in scope, whether the backend is included, jailbroken and rooted device policy | Client-side control group results plus backend API findings |
Cloud | Provider rules of engagement, platform benchmarks | Accounts and subscriptions, identity model, whether configuration review or exploitation is bought | Proven privilege escalation paths, not a configuration diff |
External network | NIST SP 800-115, PTES | IP ranges, hosting provider permissions, third-party systems excluded | Reachability and exploitation evidence per exposed service |
Internal network and Active Directory | PTES, ATT&CK | Assumed-breach starting position, domain count, trusts, hybrid identity | Full path from starting account to domain privilege, step by step |
Wireless | NIST SP 800-153, OSSTMM wireless channel | Site addresses, SSID and EAP inventory, whether client devices are in scope | Reconciled SSID inventory, proven guest-to-corporate paths |
Social engineering | PTES, OSSTMM human channel | Target population, pretexts approved, whether credentials may be captured | Per-campaign click, submit and report rates with timings |
Red team | MITRE ATT&CK, emulation plans | Objectives and crown jewels, detection posture, deconfliction process | Technique-by-technique prevented, detected or missed outcomes |
Three of these deserve a note.
Cloud is a scoping problem before it is a testing problem. The provider owns part of the stack and publishes what you may do to it. AWS names the services customers may test without approval and bans DoS, port and protocol flooding and subdomain takeovers; Azure requires no notification but binds you to its rules of engagement, which prohibit DDoS testing entirely and require you to stop and report through MSRC if you find a flaw in Microsoft's own services. Agree which side of that line your test sits on before the statement of work is signed.
Active Directory testing is usually assumed-breach, and that is correct. Starting from a standard domain user account is not a shortcut; it is the realistic starting position after a phishing email lands, and it puts the budget into the path from that account to domain privilege. Our Active Directory penetration testing guide covers what that path looks like in practice.
Wireless is the one engagement type that cannot be delivered remotely, because the medium being assessed is the radio environment itself. It also carries its own compliance driver in PCI DSS. The scope, methodology and cost detail is in our Wi-Fi penetration testing guide.
What the methodology section of a report should state
A methodology section is not a paragraph of boilerplate naming three standards. Written properly it is half a page and it answers a cold reader's questions without a follow-up call. It should state:
The standards followed and the versions. "OWASP WSTG v4.2 and OWASP ASVS 5.0.0" is checkable. "Industry best practice" is not.
The knowledge level and why. Black, grey or white box, with the credentials and documentation actually provided.
The dates and the testing window. Including out-of-hours restrictions, because a finding is a point-in-time claim.
The scope as tested, against the scope as agreed. If a host was unreachable for two days, say so.
The split between automated and manual work. The PCI SSC checklist asks for "a clear discussion of the automated and manual testing that was performed", and it is the most revealing sentence in most reports.
The severity model. Which scoring system, and whether ratings were adjusted for business context.
The limitations. Everything not attempted, and why: rate limits, availability constraints, WAF interference, third-party systems, legacy platforms excluded by agreement.
The retest terms. What was retested, when, and what remains open.
A report carrying all eight survives an auditor, an acquirer's diligence team and a customer security questionnaire. One missing the limitations section invites the reader to assume either total coverage or none, and both readings hurt you.
How buyers verify a vendor actually follows a methodology
Every vendor will say yes when asked whether they follow a methodology. These questions produce answers that differ.
"Send me the methodology section from a redacted report for an engagement like ours." The single highest-signal request in vendor selection. A firm with a real methodology has this ready. Compare what comes back against the eight items above.
"Which standards, and which versions?" Expect specific answers such as WSTG v4.2, ASVS 5.0.0, MASVS, the 2023 API Security Top 10, ATT&CK v19.x. Vagueness here predicts vagueness in the report.
"Which of your test cases are automated and which are manual?" You are looking for a vendor who can draw the line without discomfort. Everybody uses tooling. The question is what happens afterwards.
"Show me an exploitation write-up with the reproduction steps." Redacted is fine. You are checking whether findings arrive with proof or with a CVSS score and a paragraph.
"How do you scope authorisation testing?" The correct answer mentions roles, accounts at different privilege levels and object ownership. If the answer is a tool name, authorisation coverage will be thin.
"Who is testing, and what do they hold?" Named individuals and their certifications, plus whether the firm holds a company-level accreditation such as CREST. The PCI SSC supplement is explicit that qualifications "cannot be met by certifications alone" and asks about years of experience and relevant engagements, so ask about both.
"Are you organisationally independent of the systems being tested?" PCI's requirement is that the tester "must be organisationally separate from the management of the target systems", which rules out a firm testing infrastructure it installed or maintains. It is a good rule outside PCI as well.
"What is your retest policy, and how long does the window stay open?" Get this in the statement of work, not in an email.
"What happens if you find a critical issue on day one?" Immediate notification should be contractual, with a named contact and a stated timeframe.
"What will this test not tell us?" A vendor who answers this well is a vendor whose limitations section will be honest.
Put these in writing. The penetration testing RFP template formats them for a procurement process, and how to compare penetration testing quotes covers normalising the responses once they come back.
How Stingrai applies this
Stingrai is a global CREST-accredited penetration testing services company founded in Toronto, Canada in 2021, trusted by companies from startups to enterprises to meet audit requirements for SOC 2, ISO 27001, CMMC, PCI DSS and HIPAA. OSCE³, OSWE, OSEP, CREST CRT certified pentesters, who are also world-class security researchers and bug bounty hunters. Choose from fully human-led or hybrid (AI agents plus human penetration testers) engagements across web, API, mobile, AI and LLM, cloud, network, Active Directory and social engineering penetration tests and red team engagements, with findings posted to its PTaaS portal as they are confirmed, live chat with the assigned testers, Jira and Slack integration, retesting and an attestation letter with every report. Engagements are led by credentialed penetration testers from a team holding OSCE3, OSCP, OSWE, OSED, OSEP, CREST CRT, CISSP, CRTO, CRTE and eWPTX, and the firm-level CREST accreditation as a penetration testing service provider is a separate credential from the individual CREST CRT certifications testers hold. That combination is what regulated buyers in financial services, healthcare and SaaS are actually purchasing when a SOC 2, ISO 27001, PCI DSS, NYDFS or HIPAA programme asks for testing evidence. The team has published 18 CVEs and presents research at DEF CON and BSides, and Stingrai holds 5.0 out of 5.0 across 20 Clutch reviews.
In practice the methodology stack is the one described above. Web engagements run against OWASP WSTG with ASVS as the verification target; API work against the OWASP API Security Top 10 with role-separated accounts; mobile against MASVS and MASTG plus the backend; network and Active Directory work on a PTES phase structure with ATT&CK mapping where detection is in scope; cloud against the provider's own rules of engagement. Engagements are available as one-time penetration tests and as continuous testing programmes, and retest is included.
What that methodology produces is measurable. Across 55 penetration tests and 1,206 verified findings analysed in The State of Penetration Testing 2026, 51 of 55 tests, or 92.7%, surfaced at least one High or Critical issue. Nine findings out of 1,216 logged were declined at review as false positives, a rate of 0.74%, which is the number a two-stage verification step in the reporting phase is designed to hold down. Where resolution time was tracked, Critical findings closed at a median of 10.5 days.
On pricing, published packages cover one web application and its APIs: US$3,000 or US$6,800 one-time, or US$650 or US$1,275 per month on a 12-month continuous engagement, all listed on the pricing page, with a "No High or Critical Finding = Don't Pay" term on the Autonomous tier. Within that web application scope, Snipe is Stingrai's AI agent for web application penetration testing including the application's APIs, available either for autonomous web testing or working alongside penetration testers in a Hybrid web engagement. Every other service, network, cloud, Active Directory, wireless, physical and social engineering, is scoped and delivered by Stingrai's penetration testers and quoted through get a quote. Penetration testing from Stingrai supports SOC 2, ISO 27001, HIPAA, PCI DSS 4.0, NIST SP 800-53 and 800-171, DORA and NIS2 programmes by producing the technical evidence those programmes ask for.
Frequently Asked Questions
What is a penetration testing methodology?
A penetration testing methodology is the documented, repeatable process a tester follows so that the coverage of an engagement is provable rather than asserted. It sets out the phases of the test, the reference standards used for each target type, the knowledge level the tester works from, and the evidence each phase has to produce. The PCI Security Standards Council's Penetration Testing Guidance supplement treats it as a report-quality gate, asking whether the methodology is clearly stated and whether it reflects industry best practices such as OWASP and NIST.
Which penetration testing methodology should we use?
Use more than one, because each governs a different thing. Take phase structure from PTES or NIST SP 800-115, web coverage from OWASP WSTG with ASVS as the verification target, mobile coverage from OWASP MASVS and MASTG, API coverage from the OWASP API Security Top 10, and adversary behaviour from MITRE ATT&CK where the engagement is red team or adversary simulation. A vendor naming exactly one standard for every engagement type is describing a template rather than a method.
What is the difference between PTES and NIST SP 800-115?
PTES divides an engagement into seven sections, from pre-engagement interactions through intelligence gathering, threat modeling, vulnerability analysis, exploitation and post exploitation to reporting, and it deliberately carries no technical instructions. NIST SP 800-115, published September 2008, uses four phases instead, planning, discovery, attack and reporting, with a loop from attack back into discovery, and adds a rules-of-engagement template in Appendix B. PTES describes the shape of a commercial engagement more closely; SP 800-115 is the more citable reference in regulated environments.
Is OWASP a penetration testing methodology?
Partly. The OWASP Web Security Testing Guide is a testing methodology for web applications and web services, structured as numbered test cases such as WSTG-INFO-02, with v4.2 as the stable release. The OWASP Application Security Verification Standard is not a methodology but a verification target, defining pass or fail requirements across three levels, with 5.0.0 released on 30 May 2025. The OWASP Top 10 is neither: it is an awareness document listing risk categories, and a proposal that cites it as its methodology is citing the wrong document.
What are the phases of a penetration test?
In practice an engagement runs eight phases: scoping and rules of engagement, reconnaissance, threat modelling, vulnerability analysis, exploitation, post-exploitation, reporting and retest. PTES names the middle seven; NIST SP 800-115 compresses them into planning, discovery, attack and reporting, and notes that reporting runs alongside the other three rather than only at the end. Retest sits outside the classic models and is the phase auditors ask about most.
What is the difference between black box, grey box and white box testing?
Black box gives the tester no prior knowledge, which is the most realistic simulation of an external attacker and the least efficient use of a fixed budget, because discovery consumes days that would otherwise buy depth. Grey box gives partial knowledge such as credentials for each role and architecture notes, and is the default for application, network and compliance-driven testing. White box gives full knowledge including source code and configuration, reaching the most depth in authorisation logic and cryptography while telling you the least about detection. OSSTMM 3 splits the same idea into six types, separating what the tester knows from what the defender knows and expects.
Does PCI DSS require a documented penetration testing methodology?
Yes. PCI DSS Requirement 11.4 requires a defined, documented and implemented penetration testing methodology, and the PCI SSC's Penetration Testing Guidance supplement expands on what that should contain across pre-engagement, engagement and post-engagement activity. It also requires the tester to be organisationally independent, meaning separate from the management of the target systems, which prevents a firm from testing infrastructure it installed or maintains. Our PCI DSS penetration testing guide covers the requirement in detail.
How do I check that a vendor actually follows a methodology?
Ask for the methodology section from a redacted report for an engagement like yours, then check it names standards with versions, states the knowledge level and the credentials provided, distinguishes automated from manual testing, and carries a limitations section. Follow up by asking for a redacted exploitation write-up with reproduction steps, which separates firms that prove findings from firms that score them. Put both requests in the RFP rather than the sales call.
Is MITRE ATT&CK a penetration testing methodology?
No. ATT&CK is a knowledge base of adversary tactics and techniques based on real-world observations, currently at v19.2 released 28 April 2026, with 274 techniques under 15 tactics in the Enterprise matrix. It supplies the behaviour a red team emulates and the scoreboard for what your controls prevented, detected or missed, but it says nothing about scoping, rules of engagement or reporting. Use it alongside a phase model such as PTES, not instead of one.
What should the methodology section of a penetration test report contain?
It should state the standards followed with their versions, the knowledge level and the credentials or documentation actually provided, the testing dates and window, the scope as tested against the scope as agreed, the split between automated and manual work, the severity model used, the limitations and everything not attempted, and the retest terms. The PCI SSC supplement names a Statement of Methodology and a Statement of Limitations as required report sections, alongside a testing narrative and segmentation results where relevant. A report missing the limitations section leaves a reader to guess at coverage.
Related Reading
Penetration Testing Report Sample 2026: a worked example of what the methodology and findings sections should look like.
Penetration Testing Statement of Work Template 2026: the scoping and rules-of-engagement document, section by section.
How to Scope a Penetration Test: turning an asset list into a defensible scope.
How to Evaluate a Penetration Test Report: the reviewer's checklist for a report you did not commission.
The State of Penetration Testing 2026: 1,206 verified findings from 55 tests, segmented by test type.
Penetration Testing vs Vulnerability Assessment: what each compliance framework actually asks for.
Red Team vs Penetration Test vs Continuous Validation: choosing the right engagement type before choosing a methodology.
Penetration Testing Requirements by Framework 2026: which frameworks mandate testing and at what cadence.
Talk to Stingrai
Bring your asset list and we will tell you which standards apply, which phases your engagement needs, and what the deliverable will contain before you commit to anything. Book a free scoping call, get a quote, or read the published pricing.
References
The Penetration Testing Execution Standard. Main Page. http://www.pentest-standard.org/index.php/Main_Page. Defines the seven main sections of an engagement, from pre-engagement interactions to reporting, and links the separate PTES Technical Guidelines.
NIST. SP 800-115, Technical Guide to Information Security Testing and Assessment. September 2008. https://csrc.nist.gov/pubs/sp/800/115/final. Section 5.2.1 defines the four-stage penetration testing methodology; Appendix B provides a rules-of-engagement template.
OWASP. Web Security Testing Guide. Stable v4.2, released 3 December 2020. https://owasp.org/www-project-web-security-testing-guide/. Numbered test cases for web applications and web services.
OWASP. Application Security Verification Standard 5.0.0. Released 30 May 2025. https://owasp.org/www-project-application-security-verification-standard/. Pass or fail verification requirements across three priority-based levels.
OWASP. Mobile Application Security (MASVS and MASTG). https://mas.owasp.org/. Mobile control groups covering storage, cryptography, authentication, network, platform, code, resilience and privacy, with the test cases that verify them.
OWASP. API Security Top 10, 2023 edition. https://owasp.org/API-Security/editions/2023/en/0x11-t10/. The API risk classes, led by broken object level authorization.
ISECOM. OSSTMM 3, The Open Source Security Testing Methodology Manual. 2010. https://www.isecom.org/research.html. Five channels in three classes, six common test types, and the rav attack-surface metric.
MITRE. ATT&CK. Current release v19.2, 28 April 2026. https://attack.mitre.org/. Adversary tactics and techniques based on real-world observations, with Enterprise, Mobile and ICS matrices.
MITRE. Adversary Emulation Plans. https://attack.mitre.org/resources/adversary-emulation-plans/. Strategic documents and field manuals for modelling specific adversary behaviour in a red team exercise.
PCI Security Standards Council. Information Supplement: Penetration Testing Guidance, version 1.1. September 2017. https://listings.pcisecuritystandards.org/documents/Penetration-Testing-Guidance-v1_1.pdf. Methodology phases, tester qualification and organisational independence, report contents and a report-evaluation checklist.
PCI Security Standards Council. PCI DSS. https://www.pcisecuritystandards.org/standards/pci-dss/. Requirement 11.4 sets the documented penetration testing methodology obligation.
CREST. About us. https://www.crest-approved.org/. Not-for-profit accreditation body operating company accreditation and individual certification as separate credentials.
UK National Cyber Security Centre. CHECK: penetration testing. https://www.ncsc.gov.uk/information/check-penetration-testing. Standards for penetration testing that UK government, public sector and critical national infrastructure organisations can rely on.
Amazon Web Services. Penetration Testing. https://aws.amazon.com/security/penetration-testing/. Permitted services, prohibited activities, and the prior-approval requirement for command and control.
Microsoft. Penetration testing, Azure security fundamentals. https://learn.microsoft.com/en-us/azure/security/fundamentals/pen-testing. No pre-approval since 15 June 2017, with the Microsoft Cloud Unified Penetration Testing Rules of Engagement as the authoritative text.
Stingrai. The State of Penetration Testing 2026. https://www.stingrai.io/blog/state-of-penetration-testing-2026. 1,206 verified findings across 55 penetration tests, with severity, class and remediation timing.



