phishing difficultyNIST Phish Scalephishing simulation

    September 29, 2026 · 10 min read · By Fensivo Team

    Phishing simulation difficulty: the NIST Phish Scale

    Leer en español

    The conclusion in one sentence: a click rate with no declared difficulty cannot be compared

    The difficulty of a phishing simulation is how much work it takes a person to realize the email is an attack. There is a public method for rating it: the NIST Phish Scale, which combines the cues the message leaves in plain sight with how closely the email's premise matches the recipient's actual job. From that comes the conclusion that orders everything else: a click rate with no declared difficulty cannot be compared to anything. Not to last month's campaign, not to another department, not to an industry average.

    The reason is plain arithmetic, not fine methodology. A clumsy email, with an odd domain and a generic greeting, and an email that precisely mimics the invoice approval flow of a procurement team are not measuring the same thing. That holds even though both end up summarized as a percentage. If the first campaign used the clumsy email and the second used the precise one, the click rate can rise while the team improves. And it can fall while the team gets worse. When that number is the only data point that reaches the leadership committee, the noise stops being a technical detail and becomes a budget decision.

    We care about this because it is usually the missing piece. Plenty has been written about which metrics to track in a simulation program and very little about how to rate the test that produces those metrics. Without the second, the first floats.

    Human risk is managed automatically.

    Turn human risk into your first line of defense.

    Book a demo

    Free demo · 30 minutes · No commitment

    What the NIST Phish Scale is and what exactly it measures

    The Phish Scale is a NIST method, published as Technical Note 2276, "NIST Phish Scale User Guide", in November 2023, for rating the human detection difficulty of a phishing email as part of an awareness and training program. It is written for the person running the program, not for the person researching the phenomenon: its purpose is to give that person a consistent way to state how hard the email they sent actually was.

    It is worth being precise about what the method is and what it is not, because that is where the two usual mistakes happen. It does not measure the attacker's sophistication or the quality of the forgery. It does not predict how many people will fall for it, and it does not replace the click rate: it accompanies it. And it does not rate the person, it rates the email. The person remains the subject of the program's measurement. The Phish Scale is the label on the test, just as an exam carries its declared level of difficulty separately from the score each student earns.

    That distinction is what gives an external framework its value. When difficulty is declared by the vendor of the tool, it is their word against another vendor's. When it is declared by a method from a standards body, external to the category and published with its own user guide, the conversation stops being commercial and becomes verifiable.

    The two elements of difficulty: message cues and alignment with the work context

    The first element is the cues the message carries on its surface: the signs an attentive reader could notice without leaving the email. Writing errors, a sender that does not match the domain, disproportionate urgency, a link that does not correspond to the text announcing it, a request that falls outside the normal channel. The more visible cues the email carries, the easier it is to detect, and therefore the lower its difficulty.

    The second element is the one almost nobody considers when building a campaign. It is also the one that actually moves the needle: how closely the email's premise resembles what that person does all day. A payroll rejection notice does not mean the same thing to someone in human resources as it does to someone in the warehouse. A contract signature request lands differently in legal than in technical support. The same email, with exactly the same cues, is easy for someone who never receives that kind of message and hard for someone who receives it twenty times a week. Difficulty does not live in the email alone: it lives in the intersection of the email and the job.

    That intersection explains why email is still the terrain where this matters. CISA states that more than 90 percent of successful cyber attacks start with a phishing email. And what makes such an email effective is almost never a perfect forgery: it is that it arrives looking like part of the job. An attacker with a bit of prior reconnaissance does not need to write better, they need to write closer. When the simulation does not make that same effort, it is measuring how a person reacts to an attack no attacker interested in them would ever have sent.

    How it translates into an operational scale for a simulation program

    The Phish Scale gives the two axes, not a table ready to paste into a report. So the operational translation below is ours, not NIST's. It turns those two elements into a short label attached to every campaign, one that anyone understands without reading the full technical note.

    Declared levelMessage cuesPremise alignment with the jobWhat the click rate supports
    Low difficultyMany and obvious: foreign domain, poor writing, urgency without reasonGeneric, could be aimed at any employee at any companyA floor. Whoever clicks here needs immediate attention; whoever does not has proven very little
    Medium difficultyFew and subtle: credible sender, a single detectable inconsistencyFits the company or the department, but not the person's specific taskComparable across campaigns at the same level. This is the useful range for tracking month over month
    High difficultyAlmost none visible: domain, format and signature all coherentFits the role, the tool that person uses and a request they genuinely receiveA ceiling. Measuring here shows behavior under real pressure, not the ability to spot a mistake
    Not declaredUnrecordedUnrecordedNothing comparable. The number exists but has no scale, so it cannot support a conclusion

    The last row is the most important one and it describes most of the programs we see. It is not that they measure badly: it is that they measure without a label, and a result without a label cannot be reused. What the table proposes is cheap to adopt. It requires no change of tool and no change of calendar. It requires recording two more fields every time a campaign goes out, and sustaining that discipline, which is also where the cadence of the simulations themselves is decided.

    What changes in the program once difficulty is declared

    Three concrete things change, and none of them is cosmetic. The first is that the time series starts to exist: with a declared level, comparing March against September is legitimate, and without it that comparison is an act of faith. The second is that segmentation becomes honest. A department with a high click rate at high difficulty is in better shape than one with a medium click rate at low difficulty, and without the label that reading is inverted. The third is that the program can raise the bar on purpose and report it as progress instead of hiding it as decline.

    That last point draws the most resistance, and it deserves to be said plainly: a human risk management (HRM) program that only reports numbers going down has a permanent incentive to send easy emails. Declared difficulty breaks that incentive. A rising click rate stops being embarrassing, as long as the test was harder, and a falling one becomes suspicious when nothing else has changed.

    There is another reason not to settle for the bare number, and it comes from peer-reviewed evidence. The work of Ho and colleagues presented at the 2025 IEEE Symposium on Security and Privacy, and of Lain and colleagues at the 2022 edition of the same symposium, points in the same direction: completing a training course does not on its own predict a reduction in real failures. What demonstrates change is observing the behavior again. And to observe it again meaningfully, you need to know exactly how hard the previous test was. Cisco's 90-5-5 framework estimates that close to 90 percent of breaches involve a human factor. This is not a marginal concern: it is the surface where most of the risk is decided.

    Why a retest demands comparable difficulty, not the same template

    Here the framework stops being theory and becomes a design criterion. Testing someone again after they failed is the only way to know whether anything changed, but there are two ways to do it wrong. The first is resending the same template, which tests memory rather than behavior: the person recognizes the email, not the pattern. The second is sending something far easier, which guarantees a good result and means nothing.

    The right answer is a second test of comparable difficulty with a different template. Comparable does not mean similar at a glance. It means an equivalent level of visible cues and an equivalent alignment with that person's work, which is exactly what the two elements of the Phish Scale allow you to assert. That is the method's quiet contribution to a simulation program. It does not tell anyone how to write a phishing email. It gives them a language for claiming that two different emails were equally hard, and without that language the word retest is hollow.

    It is worth closing with the honest limitation. The Phish Scale is a rating method, not a statistic: it carries no effectiveness percentages and promises no outcome. Its value is that it organizes a conversation currently held in adjectives. Moving from "that was a fairly realistic email" to a declared and recorded level does not sound like much of an advance. And it is the difference between a program that accumulates data and one that accumulates anecdotes.

    At Fensivo, intelligent matching exists precisely to work on that second element: choosing, for each person, the simulation that combines their role, their prior behavior, the platforms their company actually uses and their exposed credentials, instead of sending the same email to everyone. By Fensivo's design, that combination takes effectiveness from 25 to 35 percent with random sending to 75 to 90 percent, measured by click rate and credential submission, which is another way of saying that an aligned premise raises effective difficulty. The retest is built on that base: three weeks after the failure, a simulation of the same category and sophistication but with a different template, plus a three-question conversational microlearning module passed with two correct answers out of three. That is how it works in practice for the use cases it covers today.

    Could you say today, from a record rather than from memory, how hard the last simulated email your company sent actually was?

    Sources and references

    • NIST, Technical Note 2276, "NIST Phish Scale User Guide", November 2023. csrc.nist.gov
    • CISA, "4 Things You Can Do To Keep Yourself Cyber Safe". cisa.gov
    • Cisco, "The 90-5-5 Concept: Your Key to Solving Human Risk in Cybersecurity", May 27, 2025. blogs.cisco.com
    • Ho, G. et al., "Understanding the Efficacy of Phishing Training in Practice", 2025 IEEE Symposium on Security and Privacy. ieeexplore.ieee.org
    • Lain, D., Kostiainen, K. and Čapkun, S., "Phishing in Organizations: Findings from a Large-Scale and Long-Term Study", 2022 IEEE Symposium on Security and Privacy. ieeexplore.ieee.org

    Human risk is managed automatically.

    Turn human risk into your first line of defense.

    Book a demo

    Free demo · 30 minutes · No commitment