retest protocolpassing criteriaphishing simulations

    August 13, 2026 · 7 min read · By Fensivo Team

    How to design a retest protocol and its passing criteria

    Leer en español

    A retest protocol is built from four decisions, and it pays to have them written down before you send the first simulation: when you test the person again (the window), what you repeat and what you change (the same deception category with a different template), what counts as passing (the passing criteria), and what happens while the person has not passed (the risk floor). The first step, and the one most often skipped, is setting the window: you do not retest the next day, while the lesson is still fresh, but around three weeks later, once the specific email is forgotten and what remains is what was learned, or what was not.

    We put it this bluntly because "measuring behavior change" has become a phrase half the industry repeats without almost anyone publishing how they do it. And the how is exactly what separates real validation from a dashboard that looks good. Retesting without a written protocol validates nothing: it becomes just another loose simulation that nudges the week's click rate up or down. Here is the design, step by step.

    What a retest protocol must define (the four elements)

    Before the steps, the list of what has to be decided, because an incomplete protocol fails precisely where it matters most. A well-built retest protocol defines four things, and none is optional: the time window between the failure and the new test; the category of the attack that gets repeated, along with its level of sophistication; the passing criteria, meaning which behavior counts as passing the test; and the risk floor, which is the state the person stays in until they show, more than once, that they no longer fall for it.

    Human risk is managed automatically.

    Turn human risk into your first line of defense.

    Book a demo

    Free demo · 30 minutes · No commitment

    The common mistake is treating the retest as "send another similar email and see what happens." That measures a snapshot, not a change. The difference is that each of the four elements answers a distinct question: when, what, what counts, and for how long. If any of them is missing, the result stops being interpretable. The four steps that follow are each of those decisions.

    Step 1: set the window (why three weeks, not the next day)

    Set the window at around three weeks after the failure, not the next day. The reason is simple: if you retest the person while the lesson is still fresh in memory, you measure their recall of the specific email, not their learning. The next day almost anyone recognizes the same pretext they just failed. Three weeks later, that specific memory has faded, and what you measure is whether the person internalized the warning signal or only memorized one example.

    There is peer-reviewed evidence for why this distinction matters. A study by Ho and colleagues, presented at the 2025 IEEE Symposium on Security and Privacy, found that completing training does not on its own predict a lower likelihood of falling for a real simulation. Along the same lines, the work by Lain and colleagues at the 2022 IEEE Symposium on Security and Privacy, a study spanning more than a year at a real organization, showed that phishing vulnerability is persistent and is not corrected by a single intervention.

    The practical takeaway: the timing of the test is not a logistical detail, it is what decides whether the number means anything. Too short a window inflates the result; a reasonable one makes it honest.

    Step 2: repeat the category and sophistication, change the template

    In the retest, repeat the deception category and its level of sophistication, but change the specific template. If the person failed against an executive authority pretext, the new test is also executive authority and of equivalent difficulty, but with a different sender, a different context, and different text. The logic is direct: you want to know whether they learned to recognize the type of manipulation, not whether they remember that particular email.

    Changing the category would break the measurement. If someone fell for a financial pretext and the retest arrives as a fake package notification, a good or bad result says nothing about what was taught. Keeping the sophistication is just as important: lowering the difficulty to "ensure" a pass turns the retest into a formality, and raising it abruptly measures something else. The rule is parity of category and difficulty, with a different surface. That is the control that makes before and after comparable. Here it helps to be clear about what each indicator in a simulation actually measures, a point we develop in what a retest is and why it proves behavior changed.

    Step 3: define the passing condition (what counts as passing)

    Define in writing, before sending anything, which behavior counts as passing the retest. Without this criterion, the result is left to interpretation and stops being defensible before a committee or an audit. The minimal, reasonable condition: the person does not take the dangerous action, meaning they do not hand over credentials or execute what the email asks. A stricter criterion adds that they also report the suspicious email, because reporting is the behavior that truly protects the organization.

    The passing criteria also apply to the remediation that sits between the failure and the retest. A short microlearning module, specific to the attack that was failed, with a clear passing threshold (for example, answering two of three questions correctly about the signals that were missed), makes sure the person went through the lesson before the new test. What matters is that the threshold is set in advance and is the same for everyone: a criterion decided after seeing the result is not a criterion, it is a justification. Writing it down ahead of time is what turns the retest into a measurement rather than an opinion.

    Step 4: set the risk floor until several passes in a row

    Establish that whoever fails stays in an elevated-risk state, a floor, that is not lifted by a single pass but by several consecutive passes. Here is the part almost no program formalizes. A retest passed once may be luck: the person might have been alert that day, or the pretext happened to feel familiar. Resilience is not a one-off pass, it is a pattern sustained over time.

    That is why the risk floor is designed to require consistency. As long as the person has not accumulated several tests in a row without falling, their risk score stays at the elevated level, and with it whatever safeguards the organization decides to attach to that level. This turns "they passed the test" into "they repeatedly proved they changed," which is the only thing that justifies lowering your guard. The floor also prevents the perverse effect of celebrating an improvement that was chance. How all of this translates into indicators a committee understands is something we cover in report rate, click rate and retest: what to measure.

    Why the published criterion matters more than "measuring change"

    The value of all this is not in having a retest, but in having the criterion written down and defensible. "We measure behavior change" is a claim almost every vendor in the category now makes. What few publish is the concrete protocol: how long they wait, what they repeat, what counts as passing, and how many times in a row you have to pass. That detail is the difference between a marketing promise and a methodology a CISO can audit and defend before their board.

    There is a deeper reason the criterion carries so much weight. The industry conversation began shifting from "how many people completed the training" to "how much measurable behavior changed," and that shift only makes sense if the change is measured with rules fixed in advance. A change measured with criteria adjusted afterward is not measurement, it is narrative.

    Cisco's 90-5-5 framework estimates that close to 90 percent of breaches involve a human factor, so people's behavior is where much of the risk is decided. And that behavior is best tested where the attack lands: more than 90 percent of successful cyberattacks start with a phishing email, according to CISA, so the inbox is the natural ground for a protocol that states what passing actually means.

    In practice, this is how we at Fensivo apply this protocol within a human risk management (HRM) program: every failure triggers immediate remediation and, three weeks later, a retest of the same category with a different template, a fixed passing criterion, and a risk floor that lifts only after several passes in a row. It is how we validate that the person changed and not just that they remembered an email, and you can see it in detail in our use cases.

    Before your next simulation campaign, could you show your committee the exact criterion by which you declare that an employee passed the test, or do you just have a rate that went up?

    Sources and references

    • Cisco, "The 90-5-5 Concept: Your Key to Solving Human Risk in Cybersecurity", 2025: https://blogs.cisco.com/security/the-90-5-5-concept-your-key-to-solving-human-risk-in-cybersecurity
    • CISA, "4 Things You Can Do To Keep Yourself Cyber Safe": https://www.cisa.gov/news-events/news/4-things-you-can-do-keep-yourself-cyber-safe
    • Ho, G. et al., "Understanding the Efficacy of Phishing Training in Practice", 2025 IEEE Symposium on Security and Privacy: https://ieeexplore.ieee.org/document/11023357
    • Lain, D., Kostiainen, K. and Čapkun, S., "Phishing in Organizations: Findings from a Large-Scale and Long-Term Study", 2022 IEEE Symposium on Security and Privacy: https://ieeexplore.ieee.org/document/9833766

    Human risk is managed automatically.

    Turn human risk into your first line of defense.

    Book a demo

    Free demo · 30 minutes · No commitment