
A bigger campaign isn't a better one. We explain how we apply the IASME Cyber Essentials Plus sampling pattern to phishing engagements, why 200 is the right cap, and how that translates to honest pricing.
The short version A stratified sample of around 200 users gives a statistically valid view of human risk in any organisation, regardless of headcount. We cap our phishing campaigns at that number using the same bracket pattern IASME publish for Cyber Essentials Plus device sampling. The result: bigger orgs don't pay for thousands of impractical sends, and smaller orgs aren't underscoped.
You ask three providers to quote a phishing assessment for a 5,000-employee enterprise. Two of them quote 35-50 days. One quotes 14. Why?
The two-month quotes assume you actually need to phish all 5,000 people. You don't. Real phishing engagements don't scale linearly with headcount, for three reasons that nobody explains until invoice time:
Sending 5,000 phishes from a single look-alike domain in a week burns the infrastructure. Your mail security stack flags it, your end users start warning each other, and the campaign blows by day three. The first 200 sends teach you everything the next 4,800 won't.
A sample of 200 with stratified selection (department, role, tenure) returns a 95% confidence interval at +/-7% margin. Going to 1,000 tightens that margin to about +/-3%. Going to 5,000 tightens it to under +/-2%. None of those extra precision points translate to a different remediation decision.
Modern phishing operations target high-value users (finance, IT, executive assistants) rather than broadcasting. A 5,000-user spray test simulates 1990s phishing, not 2026 phishing. Targeted spear-phishing of a curated 200 is closer to actual threat behaviour.
We didn't invent this. IASME publish a sample-size table in the Cyber Essentials Plus Test Specification: above a small population, you test a representative subset of devices rather than every machine. Auditors have used this pattern for years because it scales: an organisation with 50 laptops doesn't get the same scrutiny as one with 5,000, but the same 200-device sample produces a defensible audit conclusion in both.
We apply the same pattern to phishing populations. The brackets we use:
| Total population | Sample tested |
|---|---|
| Up to 50 | All, full population |
| 51 – 200 | Approximately one-third (ceiling N/3) |
| 201 – 500 | Approximately one-fifth (ceiling N/5) |
| 501 – 2,000 | Approximately 10% (ceiling N/10) |
| 2,000+ | 200 (cap) |
At the 2,000-user threshold the sample plateaus at 200. Above that we don't add more sends, we strengthen the stratification: representation across departments, roles, tenure bands, and geographies, so a 200-person sample of a 50,000-employee global org still tells the same story it would for a 2,000-employee one.
Every quote line on our pricing page that involves phishing surfaces the sample explicitly. If you scope 1,000 users you'll see something like:
Purple Phishing (1,000 users → sample 100, AiTM + payload) ... 13.5 days
Reporting and Debrief ............................................. 2.5 daysYou see the population (1,000), the sample (100), and the day count built from real workstreams (look-alike domain setup, AiTM proxy infrastructure, payload delivery, campaign analysis, reporting). Every line is defensible. There is no padded bench time hiding inside a fixed quote.
Sampling is right for general phishing resilience testing, where you want to measure organisational behaviour and email-stack effectiveness. It isn't right for everything:
When the assignment is to compromise a specific person (red-team objective, executive impersonation simulation), sampling doesn't apply: the target list is the scope.
If the goal is to measure click rates across the entire workforce for HR or compliance reporting, you may genuinely want full coverage. That's an awareness exercise rather than a security test.
Some sectors (NCSC CAF principle B6, certain financial regulators) expect specific population coverage. We will scope to the regulator's requirement, not the methodology default.
Multiple waves with different pretexts get separate samples. The cap is per-wave, not per-engagement, so a three-wave campaign tests up to 600 distinct users in total.
We sell methodology, not volume. A defensible 13-day phishing engagement against a 200-person sample of your 5,000-user organisation tells you what you need to know to make remediation and training decisions. The same engagement billed at 35 days against 5,000 sends tells you the same things, plus pads the invoice.
If you want a different scope, full coverage, a specific target list, multiple waves, regulator-driven population, say so during scoping and we'll quote that explicitly. What we won't do is quietly default to spraying because a bigger number is easier to sell.
One monthly plan, one front door: penetration testing, Cyber Essentials, AI security, training and advice, from £1,500 a month. No hidden costs.