Social engineering was the direct cause of 71% of all funds transfer fraud claims Coalition handled in 2025, across more than 100,000 global policyholders, and business email compromise plus funds transfer fraud together accounted for 58% of every claim the insurer saw (Coalition, 2026 Cyber Claims Report). That is the commercial case for buying a social engineering testing service in one sentence: the human layer is where the losses land, and it is the layer most security budgets test least.
A social engineering testing service is a contracted engagement in which qualified testers attempt to manipulate your people the way an attacker would, across email, voice, text and physical entry, then hand back what happened, what your controls did, and whether anyone reported it. What you are buying is tester attention and a measured result. It is not a template library and it is not a training subscription, although a mature program usually runs one of those too.
This guide is the commercial hub for the category. It compares the four assessment types, the delivery models, what a program should measure, the compliance and insurance drivers, and what it costs in 2026.
Which guide you need
This page is written for the buyer deciding what to scope and who to hire. Several deeper guides sit underneath it.
You want to | Read |
|---|---|
Understand exactly what happens inside a phishing campaign assessment | |
Understand what happens on site during a physical assessment | |
Cite the 2026 numbers on phishing, vishing and the human layer | Phishing Statistics 2026, Vishing Statistics 2026, Social Engineering Statistics 2026 |
Fix the control side rather than test it | |
Understand why MFA stopped being enough | Adversary-in-the-Middle Phishing Detection, Device Code Phishing and ClickFix |
Answer an insurer's testing questions |

What a social engineering testing service includes
Scope is defined by attacker behaviour, not by headcount. The four categories below are separately buyable and separately priced, and a competent provider will tell you which ones your risk profile actually justifies rather than selling all four.
Category | What it tests | Who it targets | Typical output |
|---|---|---|---|
Phishing simulation | Whether a researched, plausible email lure defeats your mail controls, your people and your detection stack | Whole company for a baseline, then chosen cohorts | Per-wave metric set, detection timeline, control findings |
Vishing (pretext calling) | Identity verification, callback discipline and password or MFA reset workflow | Service desk, HR, finance approvers, executive assistants | Call-by-call narrative, verification failure points, recordings where lawful |
Smishing and QR-delivered lures | Whether moving a link to an unmanaged phone routes around your controls | Mobile-heavy, field and front-line staff | Delivery and interaction rates, device policy findings |
Physical and on-site pretexting | Tailgating, reception and visitor process, badge handling, unattended access, drop-box opportunity | Reception, guarding, facilities, floor staff | Photographic evidence, entry timeline, control-layer findings |
Phishing simulation
Phishing simulation is the entry point for most buyers and the category with the widest quality spread. At the commodity floor sits a template send with a click-rate report. CISA's no-cost Cyber Hygiene Services include a Phishing Vulnerability Scanning offering that mails a mock message and reports how many users opened and clicked (CISA), which is a reasonable free baseline and a poor paid product.
An assessment-grade campaign is different in four ways: the lure is researched against your organisation rather than pulled from a library, cohorts are chosen for access and privilege rather than sprayed company-wide, the difficulty of each lure is recorded so waves are comparable, and the measurement continues past the click into what your mail gateway, identity provider, endpoint tooling and security operations team did next.
MFA-bypass-aware campaigns
The single most important change to this category since 2023 is that a campaign which stops at a fake password box will always report that MFA saved you. Microsoft's 2025 Digital Defense Report attributes 80% of MFA-bypass breaches to session-token theft via adversary-in-the-middle techniques, where a proxy captures the credential and the post-authentication session cookie together, so no second prompt is ever triggered (Microsoft).
For a buyer, this is a scoping question, not a technical one. Ask whether the provider can run a scenario that establishes whether your second factor is phishing-resistant or merely present, and ask what evidence they return. A well-run scenario answers a single defensive question, records the proof, invalidates any token obtained, and produces a recommendation about authentication policy. Our defender-side analysis of adversary-in-the-middle phishing detection covers the detection surfaces this exercises, and device code phishing and ClickFix covers the variants where there is no fake login page for a user to inspect at all.
The correct response to a successful result here is an authentication change, not another training module.
Vishing
Voice has industrialised faster than any other channel. CrowdStrike measured vishing growing 442% between the first and second halves of 2024, and reported that first-half 2025 volume alone exceeded the entire 2024 total (CrowdStrike). The reason is structural: the service desk can reset a password and re-enrol MFA, which makes it the highest-leverage human target in most organisations.
A vishing block is scoped per target rather than per mailbox, because calls are one-to-one and cannot be batched. It tests the workflow, not the person: what proof of identity is demanded, whether the agent will deviate under time pressure, whether a callback to a number of record is mandatory before a reset, and whether the deviation is logged. The full 2026 numbers sit in Vishing Statistics 2026.
Smishing and QR-delivered lures
Text and QR delivery matter because they move the interaction onto a device your controls may not manage. The finding is usually about mobile policy and conditional access rather than about the individual who scanned the code. Scope this where you have field staff, shift workers, retail floor staff or a bring-your-own-device population, and skip it where everyone works from a managed laptop on a managed network.
Physical and on-site pretexting
Physical assessment is a separate discipline with its own safety planning, authorisation paperwork and travel costs. It tests reception and visitor processes, tailgating, badge handling, unattended workstations and whether an entry can reach a live network port. Stingrai's physical security assessments service page covers the offering, and Physical Penetration Testing: What Actually Happens walks through an engagement day by day, including the authorisation letter every tester should carry.
Bundle physical with phishing and vishing when the question is "how would someone get in", and buy it standalone when the question is "does our guarding and visitor process work".
What is normally out of scope
Rule these out in writing unless separately negotiated, and treat a provider who resists as a warning sign:
Lures built on personal distress: redundancy, bereavement, illness, immigration status, discipline or debt.
Compensation themes such as bonuses, payroll errors and equity grants. They generate the highest click rates, the most complaints, and no lesson a safer premise cannot teach.
Impersonating a named employee, or a real brand, regulator or health service without written permission.
Malware execution, persistence or endpoint activity beyond an agreed telemetry marker.
Storing real passwords. A credential submission is recorded as an event, never as a captured secret.
Anything outside the agreed window, and any escalation into network exploitation or physical entry that was not scoped.
What a social engineering program actually measures
Click rate is the number every vendor leads with and the number that tells you least. It moves with the difficulty of the lure rather than the competence of your staff. NIST built the Phish Scale for exactly this problem, rating an email on observable cues and on how well its premise fits the recipient's work context, and warns that results "can create a false sense of security if click rates are analyzed on their own without understanding the phishing email's difficulty" (NIST).
Require these instead, and require them per wave and per cohort:
Measure | Why a buyer should care |
|---|---|
Delivery rate | Separates a control problem from a people problem |
Credential submission rate | The first metric that maps to actual compromise |
MFA-bypass susceptibility | Establishes whether your second factor is a barrier or a reflex |
Reporting rate | The only metric that scales defence: one report protects everyone |
Time to first report | Decides whether the security team can pull a live campaign |
Detection and response actions | Separates a report received from a report acted on |
Two of these are the ones to write into the contract. Verizon's Data Breach Investigations Report puts median time-to-click at 21 seconds against median time-to-report at 28 minutes (Verizon DBIR). Closing that gap is the entire job.
Training does move the number when it is sustained. KnowBe4's 2025 Phishing By Industry Benchmarking Report, covering 14.5 million users across 62,400 organisations and 67.7 million simulated tests, recorded the global Phish-prone Percentage falling from a 33.1% baseline to 4.1% after 12 months of consistent training, an 86% reduction (KnowBe4). The measurement detail behind each of these metrics is unpacked in what a real phishing assessment tests.
One-off assessment or continuous program
Both are legitimate purchases and they answer different questions. Buy the one that matches the decision you need to make.
One-off assessment | Continuous program | |
|---|---|---|
Question answered | If this happened on Tuesday, what would happen? | Is the human layer getting better or worse over time? |
Typical shape | One to three waves plus an optional vishing block, single report | Quarterly or monthly waves, varied difficulty, trended reporting |
Best trigger | A real incident, a new insurer question, a first baseline, an M&A diligence request | An established awareness program that has stopped producing measurable change |
Weakness | A single point-in-time reading, easy to dismiss as a bad week | Costs more, and degrades into theatre without control ownership |
Evidence value | Strong for a specific audit period or a specific board question | Strong for trend claims and renewal questionnaires |
Most organisations should start with a one-off baseline and move to a program once someone owns the control findings. A program without a named owner for gateway, identity and service-desk changes produces charts, not security.
Delivery models compared
The word "vendor" covers five very different things in this category. Comparing their prices without comparing their models is how buyers end up disappointed.
Model | Who performs the work | Strengths | Trade-offs | Published pricing |
|---|---|---|---|---|
Awareness platform, self-run | Your team, using a template library | Cheap per seat, continuous cadence, training integration, good for compliance evidence of training | Nobody with offensive experience chose the lure; measures the click and stops; no detection assessment | Yes, per seat. Microsoft prices Defender for Office 365 Plan 2, which includes cyberattack simulation training, at US$5.00 per user per month paid yearly (Microsoft) |
Offensive security firm, tester-run | Senior human testers on your engagement | Researched pretexts, cohort targeting, MFA-bypass-aware scenarios, detection and response measured, physical and vishing under one methodology | Periodic rather than continuous; costs more per campaign | Usually quoted per engagement |
Broad consultancy | A mixed bench, sometimes subcontracted | Bundles with audit and advisory; strong for large multi-jurisdiction estates | Ask who is physically doing the work; delivery quality varies by named practice | Rarely published |
MSSP or MSP bundle | Your managed provider, usually reselling a platform | Convenient, already inside your contract | Marking your own homework if the same provider runs the mail gateway being tested | Bundled, rarely itemised |
In-house, self-delivered | Your own security team | Cheapest marginal cost, deep context | Consumes the team that would otherwise be responding; independence problem when results reach the board | Not applicable |
Examples of the platform category include KnowBe4, headquartered in Clearwater, Florida and owned by Vista Equity Partners since 2023, whose platform covers phishing, vishing and smishing simulation; Proofpoint of Sunnyvale, California, whose ZenGuide product delivers risk-based awareness training; and Microsoft's Attack Simulation Training, included in Defender for Office 365 Plan 2. These are subscriptions your team operates, and they are the right purchase for cadence and training records. They are not a substitute for an assessment, and the better vendors in that category say so.
Run the platform continuously. Run the assessment periodically, to find what the platform cannot tell you.
Tooling: what to ask, and what an honest answer sounds like
Tooling is where proposals get vague. You do not need a provider to hand over their tradecraft, and you should be suspicious of one that offers to. You do need capability-level answers to five questions.
Can you run a campaign that measures more than the click? The honest answer describes purpose-built campaign infrastructure and landing pages under the provider's control, not a shared vendor domain. Stingrai's phishing campaigns service runs on a custom campaign platform built around a Gophish deployment with session-capture capability integrated, which is what makes credential-submission and MFA-interaction measurement possible rather than click-only reporting.
Can you establish whether our MFA is phishing-resistant? The answer should describe a scoped, approved scenario with token invalidation at close, and should refuse to describe operational detail in a proposal.
Do you run vishing in-house, or subcontract it? Ask for the names and certifications of the people who will be on the calls.
How is the target list handled? Transfer method, storage location, retention period, and deletion at report acceptance should all be answerable without checking.
What is in the sample report? Ask for a redacted one before signing. If it contains no reporting rate and no detection timeline, the product measures failure rather than defence.
A provider who cannot answer questions three through five in a first call is selling seats.
Compliance and insurance drivers
Two forces put this purchase on the roadmap. Neither framework below mandates a simulated attack, and any vendor telling you otherwise is misreading the standard. What they mandate is awareness and communication, and what an assessment adds is operating evidence that those controls work under adversarial conditions.
Driver | What it actually requires | What a social engineering assessment contributes |
|---|---|---|
PCI DSS v4.0.1 Requirement 12.6.3.1 | Security awareness training must include awareness of phishing and related attacks and social engineering, a future-dated requirement effective 31 March 2025 (PCI SSC) | Evidence that the training changed behaviour, with per-cohort results and a retest path |
SOC 2 Trust Services Criteria, CC2.2 | Internal communication of security responsibilities to personnel | Independent evidence that the communication is landing, plus findings on service-desk verification workflow |
ISO/IEC 27001:2022 Annex A 6.3 | Information security awareness, education and training appropriate to role | Role-targeted results by cohort rather than a single company-wide completion percentage |
Cyber insurance application forms | Testing type and frequency by category. The AXIS form 1012098 0623 asks applicants to state frequency for external network, internal network, social engineering, physical and web application testing separately | A defensible, evidenced answer in the social engineering and physical rows |
That insurance row is the one buyers most often trip on. A form asks how often you run social engineering testing, the applicant answers "annually" because the awareness platform mails a template every quarter, and the underwriting file now contains an answer the organisation cannot evidence. Our guide to pentest questions on cyber insurance underwriting forms covers the warranty and misrepresentation exposure that creates, and Coalition's own claims data explains why carriers ask: 71% of funds transfer fraud claims traced directly to social engineering, and 39% of funds transfer fraud events occurred without any confirmed email compromise at all (Coalition).
Compliance-driven demand is also what is growing the tooling market around this service. Mordor Intelligence sizes the security awareness training market at US$6.74 billion in 2026, reaching US$14.66 billion by 2031 at a 16.82% CAGR, and names cyber-insurance requirements for proof of employee training as a named growth driver (Mordor Intelligence).
For the wider compliance picture, see SOC 2 penetration testing and PCI DSS penetration testing.
What social engineering testing costs in 2026

Social engineering work is quoted per engagement rather than list-priced, because the unit of cost is tester time against your specific organisation. The useful thing to understand is the billing structure, so you can compare two quotes that look different on the surface.
Category | Billing unit | What moves the number |
|---|---|---|
Phishing simulation | Per wave, plus pretext development | Number of waves, how much bespoke pretext research each needs, number of languages and jurisdictions |
Vishing | Per target | Calls are one-to-one, so cost scales directly with target count rather than headcount |
Smishing and QR | Per wave | Delivery infrastructure and whether device-level findings are in scope |
Physical and on-site | Per site, per tester day | Site count, shift coverage, covert versus overt, travel, and whether a network-capable tester joins for a physical-to-digital pivot |
MFA-bypass scenarios | Uplift on the wave | Additional infrastructure, approvals and evidence handling |
Retest and trend reporting | Reduced rate | The design already exists, so only execution and analysis recur |
Three anchors help you sanity-check a quote. The free floor is CISA's no-cost Cyber Hygiene Services for eligible organisations. The subscription floor is platform seat pricing, where Microsoft lists Defender for Office 365 Plan 2 with cyberattack simulation training at US$5.00 per user per month paid yearly. Above that sits tester time, which is priced like any other penetration testing engagement and follows the same drivers described in Penetration Testing Cost 2026 and the average cost of a pentest in Canada.
A note on Stingrai's own pricing, because the published packages are frequently misread. The Autonomous and Hybrid packages, at US$3,000 one-time or US$450 per month and US$6,800 one-time or US$1,275 per month respectively, cover web application penetration testing. Social engineering and physical assessments are scoped modules quoted per engagement and sit in the Enterprise tier. Ask for a scoped quote with headcount, wave count, target count for vishing, and site list in hand.
Timelines
Elapsed time is driven by approvals and wave spacing, not by tester days.
Phase | Typical elapsed time |
|---|---|
Scoping, legal and HR sign-off | 1 to 2 weeks, longer where a works council must be consulted |
Reconnaissance and pretext development | 3 to 7 days per bespoke pretext |
Infrastructure build, domain ageing and quality assurance | 3 to 10 days |
Phishing wave execution | 1 to 3 days per wave, spaced 1 to 3 weeks apart |
Vishing block | 1 to 3 days |
Physical assessment, single site | 1 to 2 weeks end to end |
Analysis and reporting | 5 to 10 business days after the final activity |
A single-wave phishing engagement typically runs four to six weeks from kickoff to report. A multi-vector engagement covering phishing, vishing and one site runs eight to twelve weeks.
How to choose a social engineering testing provider
Work through these ten in order. The first four eliminate most of the field.
Establish who performs the work. Names, certifications, employment status, and whether any component is subcontracted. For vishing and physical work this matters more than for any other testing category, because judgement in the moment is the product.
Ask for a redacted sample report. Check it contains a reporting rate, a detection and response timeline, and control findings, not only a click-rate dashboard.
Confirm the pretext approval gate. Every pretext, sending domain, landing page and message body should go to a named sponsor for written approval before build. A provider who will not show you the copy in advance is running templates.
Check they can rate lure difficulty. Without a recorded difficulty rating, year-on-year comparison is meaningless. Ask which scale they use.
Test their MFA answer. Ask how they establish whether your second factor is phishing-resistant, and what they do with any session token obtained. "We invalidate at close and record proof only" is the right shape of answer.
Review the data-handling terms. Target list transfer, pseudonymisation of individual events, retention period, and deletion at report acceptance should be in the statement of work, not in a verbal assurance.
Check the ethics position. The UK National Cyber Security Centre is explicit that "blaming users for clicking on links doesn't work" (NCSC). A provider who recommends naming individuals to managers will lower your reporting rate, which is the metric you are paying to raise.
Confirm the legal scaffolding. Written authorisation, an escalation contact reachable during the window, a stop condition, and for physical work an authorisation letter carried on site. The red team rules of engagement checklist covers the clauses to demand.
Check accreditation and independence. Firm-level accreditation such as CREST tells you a third party has assessed the methodology. See CREST-accredited penetration testing companies. Independence matters if the provider also runs the mail gateway being tested.
Agree the retest path up front. What gets re-run, when, and what result counts as improvement. Findings still open at the next campaign teach the organisation that the exercise is decorative.
The pentest and red team RFP question bank has the full question set if you are running a formal procurement.
Red flags in a proposal
Warning sign | What it usually means |
|---|---|
Priced purely per mailbox | You are buying a platform seat, not tester time |
No pretext approval step | Templates, not research |
Reporting rate absent from the sample report | The product measures failure, not defence |
No detection or response timeline | Nobody is looking at your security operations |
Vague on data retention and password handling | The controls have not been designed |
Claims an automated product can assess your lobby | Physical work needs people on the ground |
Recommends naming individuals to managers | Contradicts published guidance and damages reporting rate |
Quotes vishing per mailbox rather than per target | The provider has not run vishing at scale |
Where Stingrai fits
Stingrai is a CREST-accredited penetration testing service provider at firm level, founded in 2021, with its head office in Toronto and an office in London. The team holds OSCE3, OSCP, OSWE, OSED, OSEP, CREST CRT, CISSP and CRTO among its certifications, has 18 published CVEs to its name, and holds 5.0 out of 5.0 across 19 Clutch reviews.
Social engineering engagements are delivered by senior human testers, as one-time annual assessments or as continuous multi-wave programs, whichever matches the decision you need to make. Pretext design, the judgement about when to press and when to stop, and every phone call are human work. Snipe, Stingrai's autonomous AI pentesting agent, covers web application testing and works alongside those same testers on web application engagements, so where a social engineering program includes the application a captured credential would land in, both come into scope under one methodology. Any vendor claiming an autonomous product can talk your service desk into a password reset is describing something else.
The social engineering practice covers phishing campaigns and physical security assessments. Where phishing is one initial-access vector among several rather than the whole engagement, the conversation is usually about a red team instead, and Scattered Spider identity takeover covers the account-recovery attack path that most often justifies one.
Start with a quote or the pricing page.
Frequently Asked Questions
What are social engineering testing services?
Social engineering testing services are contracted engagements in which qualified testers attempt to manipulate an organisation's people the way an attacker would, then report what happened and what the controls did. The four buyable categories are phishing simulation, vishing or pretext calling, smishing and QR-delivered lures, and physical or on-site pretexting. The output is a measured result covering delivery, credential submission, MFA-bypass susceptibility, reporting rate and how the security team responded, not a click-rate dashboard.
How do I choose between phishing simulation vendors?
Separate the two vendor types first. An awareness platform is a subscription your team operates, priced per seat, and is the right purchase for training cadence and compliance records. A tester-run assessment is an engagement in which someone with offensive experience researches your organisation, gets written approval for bespoke pretexts, and measures your controls and your security operations team. When comparing tester-run providers, ask who performs the work, request a redacted sample report, confirm the pretext approval gate, and check that the report contains a reporting rate and a detection timeline rather than click rate alone.
How much does a phishing simulation service cost?
There is no list price, because cost is driven by tester time rather than by mailbox count. The variables are the number of waves, how much bespoke pretext development each wave requires, whether MFA-bypass scenarios are in scope, the number of languages and jurisdictions, and whether vishing or physical testing is bundled in. For reference points, CISA offers a no-cost Phishing Vulnerability Scanning service under its Cyber Hygiene programme, and Microsoft lists Defender for Office 365 Plan 2 with cyberattack simulation training at US$5.00 per user per month paid yearly. Tester-run assessments are quoted per engagement.
What is a vishing assessment?
A vishing assessment is a scoped block of pretext phone calls against agreed targets, tracked by MITRE ATT&CK as T1566.004 Spearphishing Voice. It tests the workflow rather than the individual: what proof of identity is demanded, whether an agent will deviate under time pressure, whether a callback to a number of record is mandatory before a password or MFA reset, and whether the deviation is logged. It is priced per target because calls are one-to-one. The service desk is usually the highest-value population, because it is the function that can reset credentials and re-enrol MFA.
What is the best social engineering penetration testing company?
The right answer depends on which categories you need and where your sites are. Look for a firm that runs phishing, vishing and physical work under one methodology with its own senior testers rather than subcontracting the phone calls, holds firm-level accreditation such as CREST, and will show you a redacted report containing a reporting rate and a detection timeline before you sign. Stingrai delivers social engineering engagements from Toronto and London as one-time assessments or continuous programs, as a CREST-accredited penetration testing service provider founded in 2021 with a 5.0 out of 5.0 rating across 19 Clutch reviews.
Does a phishing simulation satisfy SOC 2, ISO 27001 or PCI DSS?
No framework mandates a simulated attack. PCI DSS v4.0.1 Requirement 12.6.3.1 requires that security awareness training include awareness of phishing and social engineering, effective 31 March 2025. SOC 2 criterion CC2.2 covers internal communication of security responsibilities, and ISO/IEC 27001:2022 Annex A 6.3 covers awareness, education and training. A social engineering assessment does not replace those controls. It produces operating evidence that they work under adversarial conditions, which is what auditors and underwriters increasingly ask to see. Stingrai's testing supports SOC 2, ISO 27001 and PCI DSS compliance programs by generating that evidence.
What is a good phishing simulation click rate?
There is no universal good number, because click rate moves with the difficulty of the lure rather than the competence of the staff. NIST built the Phish Scale for this reason and warns that results "can create a false sense of security if click rates are analyzed on their own without understanding the phishing email's difficulty". Judge a campaign on credential submission rate, MFA-bypass susceptibility, reporting rate and time to first report instead. Verizon's Data Breach Investigations Report puts median time-to-click at 21 seconds against median time-to-report at 28 minutes, and closing that gap is the goal.
Should we run social engineering testing once or continuously?
Both are legitimate and they answer different questions. A one-off assessment answers "if this happened on Tuesday, what would happen", and is the right purchase for a first baseline, a specific audit period, a board question or a post-incident review. A continuous program answers "is the human layer getting better or worse", and is the right purchase once an awareness program has stopped producing measurable change. Most organisations should take a one-off baseline first and move to a program once someone owns the gateway, identity and service-desk changes the baseline surfaces.
Related reading
Phishing Simulation Campaigns: What a Real Assessment Tests and What You Get
Physical Penetration Testing: What Actually Happens During an Assessment
Services: Social Engineering, Phishing Campaigns, Physical Security Assessments, Pricing
Ready to scope a social engineering engagement? Send us your headcount, the cohorts that carry the risk, and whether vishing or a site visit is in play. Get a quote or book a scoping call.



