main logo icon

Published on

August 8, 2026

|

11 min read

Phishing Simulation Campaigns: What a Real Assessment Tests and What You Get

A buyer's guide to the phishing simulation: what a tester-run campaign puts in scope, the metrics that matter beyond click rate, how pretexts get approved, the HR and privacy guardrails, what the report contains, and what drives cost.

Arafat Afzalzada

Arafat Afzalzada

Founder

Social Engineering

Summarize with AI

ChatGPTPerplexityGeminiGrokClaude

TL;DR

A phishing simulation run by penetration testers is a scoped assessment with an approved pretext, defined waves and a measured detection outcome, not an awareness platform sending templates on a schedule. Click rate alone is a vanity metric because it moves with the difficulty of the lure, which is why NIST built the Phish Scale to normalise results across campaigns. The metrics that predict real-world outcomes are credential submission rate, MFA interaction rate, reporting rate, time to first report, and whether the security team acted on the report. Campaign types worth scoping include broad credential harvest, targeted spear phishing, adversary-in-the-middle session capture, device-code and QR delivery variants, and pretext calling where vishing is in scope. Every pretext should be approved in writing by a named client sponsor before send, with themes exploiting personal distress such as bonuses, layoffs or bereavement ruled out at scoping. The UK NCSC is explicit that blaming users for clicking does not work, so a campaign designed to name and shame will damage the reporting culture it is meant to improve. Cost is driven by population size, wave count, bespoke pretext development, and whether vishing or physical access testing is bundled into the same engagement.

A phishing simulation is a controlled social engineering exercise in which testers send approved lures to a defined employee population and measure what people do, what your controls stop, and whether your security team reacts. That is a different product from an awareness platform mailing a template and reporting a click percentage.

Most buyers arrive here after a real incident, an insurer questionnaire, or an awareness renewal that produced no measurable change. This guide is written so that afterwards you can draft your own scope of work. Stingrai delivers phishing and social engineering engagements with senior human testers, because pretext design and the judgement about when to stop are human work.

What a phishing simulation is, and how it differs from an awareness tool

An awareness platform is a subscription tool your team operates: a template library, a cadence, a training redirect, an aggregate rate. It is not an assessment, because nobody with offensive experience chose the lure and nobody measured your detection stack.

A phishing campaign assessment is an engagement. Testers research the population, design a pretext specific to you, get it approved in writing, build infrastructure, run waves against agreed cohorts, and measure the chain through to whether security operations reacted. For a sense of where the commodity floor sits, CISA describes its Phishing Vulnerability Scanning service as sending a mock email and reporting "how many users opened and clicked the link" (CISA).

Dimension

Awareness platform

Tester-run assessment

Lure design

Generic template library

Researched against your org and brand

Targeting

Whole company or random slice

Cohorts chosen for access and privilege

Infrastructure

Shared vendor domains

Purpose-built, lookalike domains where approved

MFA handling

Flow stops at a fake form

Session capture and MFA interaction where scoped

Detection measured

No

Mail, endpoint, identity and SOC response

Output

Click-rate dashboard

Narrative, timeline and control findings

Question answered

Are we training people?

If this happened on Tuesday, what would happen?

Run the platform continuously; run the assessment periodically to find what it cannot tell you.

Why click rate alone is a vanity metric in a phishing simulation

Click rate moves with the difficulty of the email, not the competence of your staff. Send a broken lure and the rate drops; send one matching a live internal process and it climbs. Neither result says your organisation got safer.

NIST built the Phish Scale for exactly this. It rates an email on the observable cues that should tip a user off and on how well its premise aligns with the target's work context, producing a low, medium or high difficulty rating. NIST is blunt about the failure mode: results "can create a false sense of security if click rates are analyzed on their own without understanding the phishing email's difficulty" (NIST). Require a recorded difficulty rating for every wave, or year-on-year comparison is theatre.

Metric

What it measures

Why it matters

Delivery rate

Lures reaching the inbox rather than being blocked

Separates a control problem from a people problem

Click rate

Recipients who opened the link or attachment

Baseline exposure, comparable only at equal difficulty

Credential submission rate

Recipients who typed a username and password

The first metric mapping to actual compromise

MFA interaction rate

Recipients who approved a push or completed an authorisation

Tests whether the second factor is a barrier or a reflex

Reporting rate

Recipients who used the report button or told the service desk

The only metric that scales defence: one report protects everyone

Time to first report

Minutes from first delivery to first user report

Decides whether the SOC can pull the campaign before it spreads

SOC action and containment

Whether anyone triaged, purged the mail and secured accounts, and how fast

Separates a report received from one acted on

Reporting rate and time to first report predict whether you survive a real campaign.

Campaign types you can put in scope, and what stays out

Scope by attacker behaviour, not email count. MITRE ATT&CK identifiers give your SOC shared vocabulary.

Campaign type

ATT&CK reference

What it tests

Typical population

Broad credential harvest

T1566.002 Spearphishing Link

Baseline susceptibility, mail filtering, reporting at volume

Whole company

Targeted spear phishing

T1566.001 and T1566.002

Whether a researched, plausible lure defeats a named group

Finance, executive assistants, admins, developers

Attachment-led lure

T1566.001 Spearphishing Attachment

Attachment sandboxing, file-type policy, endpoint response

Roles taking external documents

Adversary-in-the-middle session capture

T1566.002 with credential relay

Whether MFA is phishing-resistant or merely present

Holders of sensitive SaaS access

MFA prompt pressure

T1621 MFA Request Generation

Whether users approve prompts they did not initiate

Anyone on push-based MFA

Device-code authorisation lure

Token theft via a legitimate authorisation flow

Whether staff authorise an attacker's device on a genuine vendor page

Cloud productivity suite users

QR code lure

Link delivered as a rendered code

Whether moving the link to an unmanaged phone bypasses controls

Mobile-heavy and front-line staff

Pretext calling (vishing)

T1566.004 Spearphishing Voice

Identity verification, callback discipline, reset workflow

Service desk, HR, finance approvers

Reconnaissance-only pretexting

T1598 Phishing for Information

What an attacker extracts without sending a payload

Reception, procurement, recruiters

Two deserve a note. A campaign stopping at a fake password box will always report that MFA saved you; only session capture shows whether your second factor is phishing-resistant. Device-code lures present no fake login page at all: Microsoft Threat Intelligence documented Storm-2372 directing targets to a legitimate authentication page and harvesting the tokens (Microsoft), so users trained to inspect the URL have no cue.

Rule these out in writing unless separately negotiated:

  • Lures exploiting personal distress: redundancy, bereavement, illness, immigration status, discipline or debt.

  • Compensation themes including bonus, payroll error and equity grants. They produce the highest click rates, the most complaints, and no lesson a safer premise cannot teach.

  • Impersonating a named employee, or a real brand without permission, a regulator or a health service.

  • Malware execution, persistence, or endpoint activity beyond an agreed telemetry marker.

  • Storing real passwords. A submission is recorded as an event, never as a captured secret.

  • Targeting excluded individuals, contractors outside the agreed population, or protected groups.

  • Anything outside the agreed window, or escalation into physical entry or network exploitation.

How pretexts get designed and approved before anything is sent

The pretext is the engagement. Everything else is logistics.

  1. Reconnaissance. Testers build a picture from public sources: the corporate site, job listings, press releases, conference appearances, public repositories and your technology footprint. This is what a real attacker knows for free.

  2. Premise selection. Candidates are drafted around plausible, non-distressing business events: a supplier portal migration, a tooling change, a document review, a policy acknowledgement.

  3. Difficulty rating. Each candidate is rated for cue count and premise alignment so results compare across waves.

  4. Client approval. Every pretext, sending domain, landing page and message body goes to a named sponsor for written approval before build. This is a gate, not a courtesy. If a provider will not show you the copy in advance, walk.

  5. Build and quality assurance. Infrastructure and landing pages are tested against a small internal control group to confirm rendering, tracking and safe failure.

  6. Monitored send. Waves go out with a tester watching, and pause if something unexpected happens: an unrelated incident, an outage, a bereavement in the target department.

The approval record matters beyond good manners. When an employee complains, and someone always does, HR and legal need a document showing who approved what.

Ethics, HR and privacy guardrails in social engineering penetration testing

This section decides whether the programme improves your security or quietly degrades it.

The UK National Cyber Security Centre is unambiguous. Its guidance states that "blaming users for clicking on links doesn't work" and warns that "employees who are afraid for their jobs will not report mistakes" (NCSC). A campaign designed to produce a list of names for managers works against the reporting rate you are trying to raise.

Write these into the engagement:

  • No punitive targeting. Leadership sees results aggregated by department, role or cohort. Individual data stays with a named small group.

  • Repeat clickers get support, not sanctions. If someone in accounts payable clicks every invoice lure, the finding is about your payment approval process.

  • Landing pages teach, never humiliate. Explain the cues in that email and how to report next time. No scoreboards.

  • Comms plan before and after. Leadership and the service desk know the window; everyone else learns afterwards, in a message leading with the reporting rate.

This is employee monitoring and it is regulated. The UK Information Commissioner's Office guidance on monitoring workers requires a lawful basis, says workers must normally be told before monitoring takes place, and expects a Data Protection Impact Assessment where the risk to workers' rights is high (ICO). Most programmes resolve this with a standing notice in the acceptable use policy, preserving transparency without telegraphing the date. Germany's Works Constitution Act goes further, giving the works council co-determination over "the introduction and use of technical devices designed to monitor the behaviour or performance of the employees" (BetrVG, Section 87(1) No. 6).

Data class

Retained

Visible to

Retention

Aggregate rates by cohort

Yes

Report readers

Report archive

Individual click and submission events

Pseudonymised where possible

Named security contacts

Deleted at report acceptance

Submitted passwords

No

Nobody

Not captured

Session tokens from AiTM scenarios

Proof of concept only, then invalidated

Test team, briefly

Destroyed at engagement close

Reconnaissance on individuals

Minimised to what the pretext needs

Test team

Destroyed at close

What the deliverable contains

Ask for a redacted sample before signing.

Report section

What it should contain

Executive summary

The result in plain language, the decisions leadership must make, the trend against previous campaigns

Campaign design record

Every pretext, its difficulty rating, infrastructure, cohorts and who approved each

Metric set

Delivery, click, submission, MFA interaction, reporting rate and time to first report, per wave and cohort

Per-department breakdown

Results by department, seniority band and geography, so remediation is aimed rather than sprayed

Detection and response timeline

Minute by minute from first send through first click, first report, SOC triage and containment

Control and process findings

What the mail gateway, endpoint tooling and identity provider did and did not do, plus service desk verification and approval workflows an attacker could ride without a click

Evidence pack

Lure and landing page screenshots, headers and timestamped logs, sanitised

Prioritised recommendations

Control, process and human-layer changes, each with an owner and effort estimate

Retest path

What is re-run, when, and what result counts as improvement

The timeline is what separates an assessment from a survey. It is where you learn the first report arrived in four minutes and nobody opened the mailbox for ninety.

Duration and what drives the cost of a phishing simulation service

Elapsed time is dictated by wave spacing, not tester days.

Phase

Typical elapsed time

Scoping, approvals and legal sign-off

1 to 2 weeks, longer with works council consultation

Reconnaissance and pretext development

3 to 7 days per bespoke pretext

Infrastructure build, domain ageing and QA

3 to 10 days

Wave execution

1 to 3 days per wave, spaced 1 to 3 weeks apart

Vishing block, where in scope

1 to 3 days

Analysis and reporting

5 to 10 business days after the final wave

Cost driver

Effect

Notes

Number of waves

High

Each is a discrete design, send, monitor and analysis cycle

Bespoke pretext development

High

The largest tester-time item in most engagements

Vishing

High per target

Live calling is one-to-one and cannot be batched

Physical access testing bundled in

High

A different discipline, with travel and safety planning

MFA and session-capture scenarios

Moderate to high

More infrastructure, approvals and careful handling

Population size and cohorts

Moderate

Effort scales with headcount, not linearly

Jurisdictions and languages

Moderate

Localised pretexts, separate legal review per country

Retest and trend reporting

Low to moderate

Cheaper, because the design already exists

These engagements are scoped individually, so pricing is quoted rather than listed. Start at the phishing campaign service page and request a quote via Stingrai pricing, headcount and wave count in hand.

What to prepare, and how to spot a commodity employee phishing test

Have these ready before kickoff:

  • Named approver with authority to sign off pretexts and stop the campaign, plus a deputy.

  • Escalation contact reachable by phone during the window if the SOC declares an incident.

  • Target list with exclusions applied, transferred securely.

  • Allowlisting decision. Allowlist everything and you test people with no control layer; allowlist nothing and you may test the filter instead of the people. Mature programmes run one wave each way.

  • Reporting mechanism confirmed. The report button works, routes somewhere staffed, and the SOC knows what to do with it.

  • SOC deconfliction position. Uninformed gives a true detection measurement; informed makes it a live-fire drill.

  • HR and legal sign-off, including works council consultation where it applies.

  • Post-campaign comms drafted in advance, and a named owner for control findings who can change gateway or MFA policy.

Warning sign

What it usually means

Priced purely per mailbox

You are buying a platform seat, not tester time

No pretext approval step

Templates, not research

Reporting rate absent from the sample report

The product measures failure, not defence

No detection or response timeline

Nobody is looking at your SOC

Vague on data retention and password handling

The controls are not designed

Recommends naming individuals to managers

Contradicts NCSC guidance and damages reporting

Cannot rate lure difficulty

Year-on-year comparison is meaningless

How to run the programme so behaviour actually changes

Set the goal on reporting, not clicking. Publish the reporting rate as the headline number and recognise the first reporter by role. A rising reporting rate shortens containment in a real event in a way a falling click rate never will.

Fix controls before people. If session capture worked, the answer is phishing-resistant authentication, not another training module. Human-layer improvement is the slowest lever, so pull the fast ones first. Vary lure difficulty deliberately and record it: a hard wave producing a high click rate is a finding about your control stack, not about your staff.

Test the cohorts carrying the risk. Whole-company sends are for baselines. The value is in service desk staff who can reset MFA, finance staff who can move money, engineers holding production access, and assistants holding executive inboxes.

Close the loop within two weeks. Publish results, name the control fixes with owners and dates, and state when the retest happens. Findings still open at the next campaign teach the organisation that the exercise is decorative.

Run it inside a wider programme. The social engineering hub shows how phishing sits alongside pretext calling and physical access testing, and our defences against phishing attacks guide covers the control side. For the classes most likely to defeat training, see adversary-in-the-middle phishing detection and device code phishing.

Stingrai is a CREST-accredited penetration testing service provider, founded in 2021, with offices in Toronto and London and a team holding OSCP, OSEP, OSCE3, CREST CRT, CISSP and CRTO. Phishing engagements are delivered by senior human testers. To scope one, start at the phishing campaign service page.

Frequently Asked Questions

What is a phishing simulation?

A phishing simulation is a controlled exercise in which testers send approved, realistic lures to a defined employee population and measure the outcome. A proper assessment records delivery, click, credential submission, MFA interaction, reporting rate and time to first report, and tracks whether the security team acted on the report. An awareness tool measures the click and stops.

How much does a phishing simulation cost?

There is no list price. Cost is driven by population size, number of waves, how much bespoke pretext development is required, and whether vishing or physical access testing is bundled in. Bring headcount, wave count and jurisdictions to the scoping call and request a quote through the Stingrai pricing page.

How long does a phishing campaign assessment take?

Scoping and approvals typically take one to two weeks, longer where works council consultation applies, then three to seven days per bespoke pretext. Waves take one to three days each and are spaced one to three weeks apart, with reporting adding five to ten business days afterwards. A single-wave engagement usually runs four to six weeks end to end.

What is the difference between a phishing simulation service and security awareness training?

Security awareness training is a subscription platform your team operates, mailing templates on a cadence and reporting aggregate click rates. A phishing simulation service is an engagement in which testers research your organisation, get written approval for bespoke pretexts, run defined waves against chosen cohorts, and measure your mail controls, identity controls and incident response.

What is a good phishing simulation click rate?

There is no universal good number, because click rate moves with the difficulty of the lure rather than the competence of your staff. NIST built the Phish Scale for this reason and warns that results "can create a false sense of security if click rates are analyzed on their own without understanding the phishing email's difficulty" (NIST). Judge a campaign on credential submission rate, reporting rate and time to first report.

Can you run an employee phishing test without telling employees?

You should avoid announcing the specific date, but you should not run one with no notice at all. The UK Information Commissioner's Office guidance on monitoring workers requires a lawful basis, says workers must normally be told before monitoring takes place, and expects a Data Protection Impact Assessment where the risk is high (ICO). Most programmes resolve this with a standing notice in the acceptable use policy.

Simulated phishing is lawful when run as transparent, proportionate employee monitoring with an identified lawful basis, minimised data collection and a documented impact assessment where the risk warrants one. Passwords should never be captured or stored, and retention periods belong in the statement of work. Germany's Works Constitution Act also gives the works council co-determination over technical devices designed to monitor employee behaviour or performance (BetrVG, Section 87(1) No. 6).

Does social engineering penetration testing include phone calls or vishing?

It can, and pretext calling is often the highest-value component because it targets the service desk, the function that can reset passwords and re-enrol MFA. MITRE ATT&CK tracks it as T1566.004 Spearphishing Voice, and a scoped vishing block tests identity verification, callback discipline and reset workflow. It is priced per target, because calls are one-to-one.

0 views

0

X

Related reading

Physical Penetration Testing: What Actually Happens During an Assessment
Social Engineering

Physical Penetration Testing: What Actually Happens During an Assessment

What happens during a physical penetration test: scope, covert vs overt, objectives, the authorization letter testers carry, findings, deliverables and cost.

11 min read

Red Team Rules of Engagement: What to Demand Before You Sign
Network SecuritySocial Engineering

Red Team Rules of Engagement: What to Demand Before You Sign

A buyer's checklist of the red team rules of engagement clauses to demand before you sign, plus a clear answer on whether testing can break production.

17 min read

Scattered Spider Identity Takeover: A Buyer's Guide to Account Recovery Red Teaming
Social EngineeringAdvisories

Scattered Spider Identity Takeover: A Buyer's Guide to Account Recovery Red Teaming

How to buy a red team that tests your account-recovery and help-desk workflows against Scattered Spider social engineering, plus a resilience checklist.

11 min read

Contents

X