Does phishing simulation training actually work?
By Elias Lankinen11 min read
I have strong, well-dated sourcing and three verified direct image URLs. Writing the post now.
The $650 That Wasn't There
In mid-December 2020, roughly 500 GoDaddy employees opened an email announcing a "$650 one-time Holiday bonus." All they had to do was fill in some details by Friday. This was a hard year: the pandemic was raging, and the company had already told staff there would be no bonus. Then the money appeared anyway. Some people filled out the form. Two days later, according to Engadget, a second email arrived from the chief security officer: "You're getting this email because you failed our recent phishing test." There was no bonus. It was a simulation, and they had flunked it. The backlash was immediate and brutal. GoDaddy apologized. But the episode captured, in miniature, the strange bargain at the heart of a multi-billion-dollar industry: to teach employees not to be deceived, companies deceive them, then grade them on the result. The premise is that this makes people harder to fool. A remarkable amount of recent evidence suggests it mostly does not.
Phishing, for the uninitiated, is the practice of sending fraudulent messages that impersonate someone trustworthy to trick a recipient into clicking a malicious link, opening a booby-trapped attachment, or handing over a password. It is not a fringe threat. Verizon's 2025 Data Breach Investigations Report, the industry's most-cited annual autopsy of real breaches, finds the "human element" involved in roughly 60 percent of them. The standard corporate response is phishing simulation training: security teams send their own fake phishing emails, see who clicks, and route the clickers into remedial lessons. The global market for this kind of security awareness training runs into the billions. The question almost nobody asked rigorously until recently is whether it works.
The biggest experiments say: barely
The most important piece of evidence arrived in 2025. Researchers led by Grant Ho, along with Ariana Mirian, Stefan Savage, and Geoffrey Voelker, ran what is likely the largest controlled study of anti-phishing training ever conducted, and presented it at the IEEE Symposium on Security and Privacy. Over eight months they sent ten simulated phishing campaigns to more than 19,500 employees of UC San Diego Health, a large hospital system, and randomly assigned people to different conditions so they could measure cause and effect rather than just correlation. The findings, reported by UC San Diego, were bleak for the industry. Annual cybersecurity training, the mandatory once-a-year module most office workers know and dread, showed no significant relationship with whether someone later failed a phishing test. "Embedded" training, the lesson served up immediately after someone clicks a simulated lure, reduced the likelihood of clicking by about two percentage points. Not two percent better than before, but a two-point absolute difference, which the authors treated as marginal. Their conclusion was blunt: anti-phishing training programs "in their current and commonly deployed forms are unlikely to offer significant practical value in reducing the risk of successful phishing attacks." The numbers underneath are worth sitting with. In the first month of the study, about 10 percent of employees clicked a phishing link. By the eighth month, more than half had clicked at least one. This was a workforce being actively, repeatedly trained, and cumulative exposure kept finding new victims. Susceptibility also depended enormously on the bait: a fake "password update" notice fooled 1.82 percent of recipients, while a message about a change to the vacation policy pulled in 30.8 percent. The lure mattered far more than the training. This was not a lone result. Three years earlier, Daniele Lain, Kari Kostiainen, and Srdjan Capkun of ETH Zurich published Phishing in Organizations, a 15-month study of roughly 14,000 employees at a partner company, also at IEEE S&P. They reached a conclusion that should have rattled the industry more than it did: embedded training delivered during simulated phishing exercises "does not make employees more resilient to phishing." In some of their data it appeared to make things slightly worse, possibly because being handed a lesson at the moment of failure teaches people less about phishing than about resenting the security team. Two large, independent, real-world experiments, using the gold-standard method of random assignment, both landed in roughly the same place. That convergence is why this stopped being a niche academic quarrel and started showing up in the mainstream press.
The click rate is the wrong scoreboard
Here is the most common misconception, and it is baked into how most programs are sold: that the goal is to drive the click rate to zero, and that a falling click rate proves the training is working. Both halves of that are shaky. Verizon's 2025 report notes that the median time for someone to click a phishing link after opening the email is under 60 seconds. A single successful click can be enough for an attacker to steal a credential or plant malware. If your defense is a workforce of thousands and your standard is that not one of them clicks within a minute, ever, you have designed a system guaranteed to fail. Real attackers only need one person, one time. A program that cuts clicking from 10 percent to 8 percent has changed almost nothing about that arithmetic. The falling-click-rate story is also easy to manufacture. Send easy simulations and the click rate drops. Warn people that "phishing season" is coming and it drops. Let employees learn the visual tells of your security vendor's templates, which look subtly different from real mail, and they will ace your tests while staying just as vulnerable to a genuinely novel attack. A low simulation failure rate can measure how well people recognize your fakes rather than how well they resist the real thing. This is where the popular idea of the "human firewall," the notion that a well-drilled workforce can be the organization's primary line of defense, starts to look less like a strategy and more like a way of shifting blame. You would not run a hospital whose infection-control plan was "we told the staff to be careful." The Ho team's own recommendation was to lean harder on technical controls that do not depend on any individual making the right split-second decision: multi-factor authentication, password managers that only autofill on legitimate domains, and phishing-resistant login methods.
The lessons nobody reads
If you want to understand why the training moves the needle so little, look at what people actually do when it appears. In the UC San Diego study, three-quarters of employees who were shown embedded training after clicking spent one minute or less on the material. About a third closed the page immediately, without engaging at all. The lesson was delivered. It was simply not received. This is not surprising. The training arrives at the worst possible moment, when a busy person has just been told they did something wrong, and it competes with every other demand on their attention. The implicit model, that a person who clicks is missing a piece of knowledge, and that supplying that knowledge on the spot fixes the gap, does not match how the mistakes happen. People rarely click because they do not know phishing exists. They click because they were rushed, distracted, or because the message was genuinely good. There is a subtler point here about design. A study from the University of South Florida's Muma College of Business, led by Dezhi Yin and Matthew Mullarkey with colleagues at other institutions and published in MIS Quarterly in November 2025, ran three experiments with more than 12,000 participants. It compared the standard approach, instant feedback delivered only to the people who failed, against delayed feedback delivered to everyone, whether they clicked or not. The inclusive, time-delayed version produced better recognition of scams that lasted for months. Training only the people who failed, at the moment they failed, was the weaker option, in part because it put people on the defensive exactly when they were least able to learn. The finding directly contradicts the "embedded" model that most commercial platforms are built around.
So when does it work?
It would be too tidy to say training never helps. The honest picture is that some approaches, run in some ways, produce measurable improvement, and that the industry's standard product is usually not one of them. The signal that recurs across the more optimistic evidence is not clicking, but reporting: whether employees flag suspicious messages so the security team can act. The ETH Zurich study, for all its skepticism about training, found real value in giving employees a simple button to report suspected phishing, then treating the workforce as a distributed sensor network. Reports came in fast, participation held up over time, and the crowd surfaced new campaigns quickly. This reframes the employee from a lock that must never fail into an alarm that can ring. Alarms are allowed to be imperfect, because you only need one to go off. Vendors have seized on this. Companies like Hoxhunt and KnowBe4 publish striking figures: KnowBe4 claims an 86 percent drop in click rates over a year of its training, and Hoxhunt reports pushing simulated-threat reporting from around 10 percent to 60 percent by running roughly 40 short, personalized simulations per user per year instead of the usual four. Their argument, made explicitly in response to a Wall Street Journal piece that cited the UC San Diego research, is that the studies condemning training tested exactly the lazy, annual, one-size-fits-all version that everyone agrees is useless, and that continuous, adaptive, skill-building programs are a different animal. There is something to this. Frequent, short, personalized practice matching the difficulty to the person genuinely looks better in the data than an annual slideshow, and Verizon's 2025 report notes that people trained more recently do report phishing at meaningfully higher rates. But two cautions apply. First, this evidence comes overwhelmingly from the vendors selling the product, using their own metrics, without the random assignment and independent control groups that made the pessimistic studies credible. Reporting rates and self-selected customer case studies are not the same currency as a randomized trial. Second, even the friendliest reading shows the mechanism that works is reporting and rapid detection, not the fantasy of an unclickable workforce.
The part that isn't measured on a dashboard
GoDaddy's fake bonus was not just a public-relations mistake. It exposed a cost that click rates never capture: what deceiving your own staff does to the relationship between them and the people meant to protect them. Research is beginning to quantify this. Work out of the University of Sussex has found that employees subjected to deceptive security testing report lower trust in leadership, and start to wonder whether their employer is on their side or simply waiting for them to slip. This is not a soft concern. If the security team is experienced as a gotcha machine, the rational employee response to a real suspicious email is to say nothing, because reporting invites scrutiny and clicking might get you punished. That directly undermines the one thing, reporting, that the evidence actually supports. Punitive simulations can suppress the behavior you most want. The ethics have become a live research question in their own right. A 2025 vignette experiment presented at the USEC workshop asked people what makes a phishing simulation acceptable or not, and found that emotionally manipulative lures, dangling bonuses, health benefits, layoffs, the exact register GoDaddy struck, are precisely the ones people judge as crossing a line. The uncomfortable irony is that those cruel lures are also the effective ones, because real attackers use them. A simulation that only sends obvious, low-stakes fakes is honest but unrealistic; a simulation that faithfully mimics a real attacker's emotional manipulation is realistic but, many employees feel, a betrayal. There is no clean way out of that tension, and most programs resolve it by not thinking about it.
What to watch
The center of gravity is shifting, and it is worth watching where it lands. The clearest movement is away from training as a preventive control and toward technical measures that do not ask a human to be perfect at 4:55 on a Friday. Phishing-resistant authentication, the passkeys and hardware security keys now built into major platforms, can make a stolen password close to useless, which changes the value of a successful phish from "catastrophe" to "nuisance." Where that becomes standard, the entire premise of phishing simulation, that the click is the disaster to be prevented, weakens. The second thing to watch is generative AI, working both ends of the problem. Attackers now use large language models to write flawless, personalized lures at scale, erasing the clumsy grammar that training taught a generation of workers to look for. The visual and linguistic tells are disappearing. At the same time, defenders are using the same models to filter mail and to spot anomalies faster than any human sensor network. If both trends hold, the human in the middle, the employee being trained to eyeball a message and judge it, may be squeezed out of relevance from both sides. Which returns us to the original bargain. The reason to run a phishing simulation was always that the human was the weakest link and could be strengthened. The best current evidence says the strengthening is small, easily overstated, sometimes counterproductive, and purchased at a real cost in trust. That does not make the whole enterprise worthless. It makes it a modest tool oversold as a solution, and worth keeping only in the humble, reporting-focused, non-punitive form the evidence actually supports. The question is no longer whether we can train people out of being human. It is whether we ever needed them not to be.
Sources
- Engadget, "GoDaddy phishing 'test' teased employees with a fake holiday bonus", 2020.
- Verizon, 2025 Data Breach Investigations Report, 2025.
- UC San Diego Today, "Cybersecurity Training Programs Don't Prevent Employees from Falling for Phishing Scams", 2025.
- Grant Ho, Ariana Mirian, Stefan Savage, Geoffrey M. Voelker et al., "Understanding the Efficacy of Phishing Training in Practice", IEEE Symposium on Security and Privacy, 2025.
- Daniele Lain, Kari Kostiainen, Srdjan Capkun, "Phishing in Organizations: Findings from a Large-Scale and Long-Term Study", IEEE Symposium on Security and Privacy, 2022.
- EurekAlert / University of South Florida, "USF study finds smarter way to train employees to thwart phishing scams" (Yin, Mullarkey, de Vreede, Limayem, MIS Quarterly), 2025.
- KnowBe4, "KnowBe4 Report Reveals Security Training Reduces Global Phishing Click Rates by 86%", 2023.
- Hoxhunt, "Why the Wall Street Journal Got It Wrong: Phishing Training Works When Done Right", 2025.
- USEC 2025 (NDSS Symposium), "What Makes Phishing Simulation Campaigns (Un)Acceptable? A Vignette Experiment", 2025.