A phishing simulation is a controlled social engineering exercise in which testers send approved lures to a defined employee population and measure what people do, what your controls stop, and whether your security team reacts. That is a different product from an awareness platform mailing a template and reporting a click percentage.
Most buyers arrive here after a real incident, an insurer questionnaire, or an awareness renewal that produced no measurable change. This guide is written so that afterwards you can draft your own scope of work. Stingrai delivers phishing and social engineering engagements with senior human testers, because pretext design and the judgement about when to stop are human work.
What a phishing simulation is, and how it differs from an awareness tool
An awareness platform is a subscription tool your team operates: a template library, a cadence, a training redirect, an aggregate rate. It is not an assessment, because nobody with offensive experience chose the lure and nobody measured your detection stack.
A phishing campaign assessment is an engagement. Testers research the population, design a pretext specific to you, get it approved in writing, build infrastructure, run waves against agreed cohorts, and measure the chain through to whether security operations reacted. For a sense of where the commodity floor sits, CISA describes its Phishing Vulnerability Scanning service as sending a mock email and reporting "how many users opened and clicked the link" (CISA).
Dimension | Awareness platform | Tester-run assessment |
|---|---|---|
Lure design | Generic template library | Researched against your org and brand |
Targeting | Whole company or random slice | Cohorts chosen for access and privilege |
Infrastructure | Shared vendor domains | Purpose-built, lookalike domains where approved |
MFA handling | Flow stops at a fake form | Session capture and MFA interaction where scoped |
Detection measured | No | Mail, endpoint, identity and SOC response |
Output | Click-rate dashboard | Narrative, timeline and control findings |
Question answered | Are we training people? | If this happened on Tuesday, what would happen? |
Run the platform continuously; run the assessment periodically to find what it cannot tell you.
Why click rate alone is a vanity metric in a phishing simulation
Click rate moves with the difficulty of the email, not the competence of your staff. Send a broken lure and the rate drops; send one matching a live internal process and it climbs. Neither result says your organisation got safer.
NIST built the Phish Scale for exactly this. It rates an email on the observable cues that should tip a user off and on how well its premise aligns with the target's work context, producing a low, medium or high difficulty rating. NIST is blunt about the failure mode: results "can create a false sense of security if click rates are analyzed on their own without understanding the phishing email's difficulty" (NIST). Require a recorded difficulty rating for every wave, or year-on-year comparison is theatre.
Metric | What it measures | Why it matters |
|---|---|---|
Delivery rate | Lures reaching the inbox rather than being blocked | Separates a control problem from a people problem |
Click rate | Recipients who opened the link or attachment | Baseline exposure, comparable only at equal difficulty |
Credential submission rate | Recipients who typed a username and password | The first metric mapping to actual compromise |
MFA interaction rate | Recipients who approved a push or completed an authorisation | Tests whether the second factor is a barrier or a reflex |
Reporting rate | Recipients who used the report button or told the service desk | The only metric that scales defence: one report protects everyone |
Time to first report | Minutes from first delivery to first user report | Decides whether the SOC can pull the campaign before it spreads |
SOC action and containment | Whether anyone triaged, purged the mail and secured accounts, and how fast | Separates a report received from one acted on |
Reporting rate and time to first report predict whether you survive a real campaign.
Campaign types you can put in scope, and what stays out
Scope by attacker behaviour, not email count. MITRE ATT&CK identifiers give your SOC shared vocabulary.
Campaign type | ATT&CK reference | What it tests | Typical population |
|---|---|---|---|
Broad credential harvest | T1566.002 Spearphishing Link | Baseline susceptibility, mail filtering, reporting at volume | Whole company |
Targeted spear phishing | T1566.001 and T1566.002 | Whether a researched, plausible lure defeats a named group | Finance, executive assistants, admins, developers |
Attachment-led lure | T1566.001 Spearphishing Attachment | Attachment sandboxing, file-type policy, endpoint response | Roles taking external documents |
Adversary-in-the-middle session capture | T1566.002 with credential relay | Whether MFA is phishing-resistant or merely present | Holders of sensitive SaaS access |
MFA prompt pressure | T1621 MFA Request Generation | Whether users approve prompts they did not initiate | Anyone on push-based MFA |
Device-code authorisation lure | Token theft via a legitimate authorisation flow | Whether staff authorise an attacker's device on a genuine vendor page | Cloud productivity suite users |
QR code lure | Link delivered as a rendered code | Whether moving the link to an unmanaged phone bypasses controls | Mobile-heavy and front-line staff |
Pretext calling (vishing) | T1566.004 Spearphishing Voice | Identity verification, callback discipline, reset workflow | Service desk, HR, finance approvers |
Reconnaissance-only pretexting | T1598 Phishing for Information | What an attacker extracts without sending a payload | Reception, procurement, recruiters |
Two deserve a note. A campaign stopping at a fake password box will always report that MFA saved you; only session capture shows whether your second factor is phishing-resistant. Device-code lures present no fake login page at all: Microsoft Threat Intelligence documented Storm-2372 directing targets to a legitimate authentication page and harvesting the tokens (Microsoft), so users trained to inspect the URL have no cue.
Rule these out in writing unless separately negotiated:
Lures exploiting personal distress: redundancy, bereavement, illness, immigration status, discipline or debt.
Compensation themes including bonus, payroll error and equity grants. They produce the highest click rates, the most complaints, and no lesson a safer premise cannot teach.
Impersonating a named employee, or a real brand without permission, a regulator or a health service.
Malware execution, persistence, or endpoint activity beyond an agreed telemetry marker.
Storing real passwords. A submission is recorded as an event, never as a captured secret.
Targeting excluded individuals, contractors outside the agreed population, or protected groups.
Anything outside the agreed window, or escalation into physical entry or network exploitation.
How pretexts get designed and approved before anything is sent
The pretext is the engagement. Everything else is logistics.
Reconnaissance. Testers build a picture from public sources: the corporate site, job listings, press releases, conference appearances, public repositories and your technology footprint. This is what a real attacker knows for free.
Premise selection. Candidates are drafted around plausible, non-distressing business events: a supplier portal migration, a tooling change, a document review, a policy acknowledgement.
Difficulty rating. Each candidate is rated for cue count and premise alignment so results compare across waves.
Client approval. Every pretext, sending domain, landing page and message body goes to a named sponsor for written approval before build. This is a gate, not a courtesy. If a provider will not show you the copy in advance, walk.
Build and quality assurance. Infrastructure and landing pages are tested against a small internal control group to confirm rendering, tracking and safe failure.
Monitored send. Waves go out with a tester watching, and pause if something unexpected happens: an unrelated incident, an outage, a bereavement in the target department.
The approval record matters beyond good manners. When an employee complains, and someone always does, HR and legal need a document showing who approved what.
Ethics, HR and privacy guardrails in social engineering penetration testing
This section decides whether the programme improves your security or quietly degrades it.
The UK National Cyber Security Centre is unambiguous. Its guidance states that "blaming users for clicking on links doesn't work" and warns that "employees who are afraid for their jobs will not report mistakes" (NCSC). A campaign designed to produce a list of names for managers works against the reporting rate you are trying to raise.
Write these into the engagement:
No punitive targeting. Leadership sees results aggregated by department, role or cohort. Individual data stays with a named small group.
Repeat clickers get support, not sanctions. If someone in accounts payable clicks every invoice lure, the finding is about your payment approval process.
Landing pages teach, never humiliate. Explain the cues in that email and how to report next time. No scoreboards.
Comms plan before and after. Leadership and the service desk know the window; everyone else learns afterwards, in a message leading with the reporting rate.
This is employee monitoring and it is regulated. The UK Information Commissioner's Office guidance on monitoring workers requires a lawful basis, says workers must normally be told before monitoring takes place, and expects a Data Protection Impact Assessment where the risk to workers' rights is high (ICO). Most programmes resolve this with a standing notice in the acceptable use policy, preserving transparency without telegraphing the date. Germany's Works Constitution Act goes further, giving the works council co-determination over "the introduction and use of technical devices designed to monitor the behaviour or performance of the employees" (BetrVG, Section 87(1) No. 6).
Data class | Retained | Visible to | Retention |
|---|---|---|---|
Aggregate rates by cohort | Yes | Report readers | Report archive |
Individual click and submission events | Pseudonymised where possible | Named security contacts | Deleted at report acceptance |
Submitted passwords | No | Nobody | Not captured |
Session tokens from AiTM scenarios | Proof of concept only, then invalidated | Test team, briefly | Destroyed at engagement close |
Reconnaissance on individuals | Minimised to what the pretext needs | Test team | Destroyed at close |
What the deliverable contains
Ask for a redacted sample before signing.
Report section | What it should contain |
|---|---|
Executive summary | The result in plain language, the decisions leadership must make, the trend against previous campaigns |
Campaign design record | Every pretext, its difficulty rating, infrastructure, cohorts and who approved each |
Metric set | Delivery, click, submission, MFA interaction, reporting rate and time to first report, per wave and cohort |
Per-department breakdown | Results by department, seniority band and geography, so remediation is aimed rather than sprayed |
Detection and response timeline | Minute by minute from first send through first click, first report, SOC triage and containment |
Control and process findings | What the mail gateway, endpoint tooling and identity provider did and did not do, plus service desk verification and approval workflows an attacker could ride without a click |
Evidence pack | Lure and landing page screenshots, headers and timestamped logs, sanitised |
Prioritised recommendations | Control, process and human-layer changes, each with an owner and effort estimate |
Retest path | What is re-run, when, and what result counts as improvement |
The timeline is what separates an assessment from a survey. It is where you learn the first report arrived in four minutes and nobody opened the mailbox for ninety.
Duration and what drives the cost of a phishing simulation service
Elapsed time is dictated by wave spacing, not tester days.
Phase | Typical elapsed time |
|---|---|
Scoping, approvals and legal sign-off | 1 to 2 weeks, longer with works council consultation |
Reconnaissance and pretext development | 3 to 7 days per bespoke pretext |
Infrastructure build, domain ageing and QA | 3 to 10 days |
Wave execution | 1 to 3 days per wave, spaced 1 to 3 weeks apart |
Vishing block, where in scope | 1 to 3 days |
Analysis and reporting | 5 to 10 business days after the final wave |
Cost driver | Effect | Notes |
|---|---|---|
Number of waves | High | Each is a discrete design, send, monitor and analysis cycle |
Bespoke pretext development | High | The largest tester-time item in most engagements |
Vishing | High per target | Live calling is one-to-one and cannot be batched |
Physical access testing bundled in | High | A different discipline, with travel and safety planning |
MFA and session-capture scenarios | Moderate to high | More infrastructure, approvals and careful handling |
Population size and cohorts | Moderate | Effort scales with headcount, not linearly |
Jurisdictions and languages | Moderate | Localised pretexts, separate legal review per country |
Retest and trend reporting | Low to moderate | Cheaper, because the design already exists |
These engagements are scoped individually, so pricing is quoted rather than listed. Start at the phishing campaign service page and request a quote via Stingrai pricing, headcount and wave count in hand.
What to prepare, and how to spot a commodity employee phishing test
Have these ready before kickoff:
Named approver with authority to sign off pretexts and stop the campaign, plus a deputy.
Escalation contact reachable by phone during the window if the SOC declares an incident.
Target list with exclusions applied, transferred securely.
Allowlisting decision. Allowlist everything and you test people with no control layer; allowlist nothing and you may test the filter instead of the people. Mature programmes run one wave each way.
Reporting mechanism confirmed. The report button works, routes somewhere staffed, and the SOC knows what to do with it.
SOC deconfliction position. Uninformed gives a true detection measurement; informed makes it a live-fire drill.
HR and legal sign-off, including works council consultation where it applies.
Post-campaign comms drafted in advance, and a named owner for control findings who can change gateway or MFA policy.
Warning sign | What it usually means |
|---|---|
Priced purely per mailbox | You are buying a platform seat, not tester time |
No pretext approval step | Templates, not research |
Reporting rate absent from the sample report | The product measures failure, not defence |
No detection or response timeline | Nobody is looking at your SOC |
Vague on data retention and password handling | The controls are not designed |
Recommends naming individuals to managers | Contradicts NCSC guidance and damages reporting |
Cannot rate lure difficulty | Year-on-year comparison is meaningless |
How to run the programme so behaviour actually changes
Set the goal on reporting, not clicking. Publish the reporting rate as the headline number and recognise the first reporter by role. A rising reporting rate shortens containment in a real event in a way a falling click rate never will.
Fix controls before people. If session capture worked, the answer is phishing-resistant authentication, not another training module. Human-layer improvement is the slowest lever, so pull the fast ones first. Vary lure difficulty deliberately and record it: a hard wave producing a high click rate is a finding about your control stack, not about your staff.
Test the cohorts carrying the risk. Whole-company sends are for baselines. The value is in service desk staff who can reset MFA, finance staff who can move money, engineers holding production access, and assistants holding executive inboxes.
Close the loop within two weeks. Publish results, name the control fixes with owners and dates, and state when the retest happens. Findings still open at the next campaign teach the organisation that the exercise is decorative.
Run it inside a wider programme. The social engineering hub shows how phishing sits alongside pretext calling and physical access testing, and our defences against phishing attacks guide covers the control side. For the classes most likely to defeat training, see adversary-in-the-middle phishing detection and device code phishing.
Stingrai is a CREST-accredited penetration testing service provider, founded in 2021, with offices in Toronto and London and a team holding OSCP, OSEP, OSCE3, CREST CRT, CISSP and CRTO. Phishing engagements are delivered by senior human testers. To scope one, start at the phishing campaign service page.
Frequently Asked Questions
What is a phishing simulation?
A phishing simulation is a controlled exercise in which testers send approved, realistic lures to a defined employee population and measure the outcome. A proper assessment records delivery, click, credential submission, MFA interaction, reporting rate and time to first report, and tracks whether the security team acted on the report. An awareness tool measures the click and stops.
How much does a phishing simulation cost?
There is no list price. Cost is driven by population size, number of waves, how much bespoke pretext development is required, and whether vishing or physical access testing is bundled in. Bring headcount, wave count and jurisdictions to the scoping call and request a quote through the Stingrai pricing page.
How long does a phishing campaign assessment take?
Scoping and approvals typically take one to two weeks, longer where works council consultation applies, then three to seven days per bespoke pretext. Waves take one to three days each and are spaced one to three weeks apart, with reporting adding five to ten business days afterwards. A single-wave engagement usually runs four to six weeks end to end.
What is the difference between a phishing simulation service and security awareness training?
Security awareness training is a subscription platform your team operates, mailing templates on a cadence and reporting aggregate click rates. A phishing simulation service is an engagement in which testers research your organisation, get written approval for bespoke pretexts, run defined waves against chosen cohorts, and measure your mail controls, identity controls and incident response.
What is a good phishing simulation click rate?
There is no universal good number, because click rate moves with the difficulty of the lure rather than the competence of your staff. NIST built the Phish Scale for this reason and warns that results "can create a false sense of security if click rates are analyzed on their own without understanding the phishing email's difficulty" (NIST). Judge a campaign on credential submission rate, reporting rate and time to first report.
Can you run an employee phishing test without telling employees?
You should avoid announcing the specific date, but you should not run one with no notice at all. The UK Information Commissioner's Office guidance on monitoring workers requires a lawful basis, says workers must normally be told before monitoring takes place, and expects a Data Protection Impact Assessment where the risk is high (ICO). Most programmes resolve this with a standing notice in the acceptable use policy.
Is phishing simulation legal under GDPR in the UK and EU?
Simulated phishing is lawful when run as transparent, proportionate employee monitoring with an identified lawful basis, minimised data collection and a documented impact assessment where the risk warrants one. Passwords should never be captured or stored, and retention periods belong in the statement of work. Germany's Works Constitution Act also gives the works council co-determination over technical devices designed to monitor employee behaviour or performance (BetrVG, Section 87(1) No. 6).
Does social engineering penetration testing include phone calls or vishing?
It can, and pretext calling is often the highest-value component because it targets the service desk, the function that can reset passwords and re-enrol MFA. MITRE ATT&CK tracks it as T1566.004 Spearphishing Voice, and a scoped vishing block tests identity verification, callback discipline and reset workflow. It is priced per target, because calls are one-to-one.



