human risk managementsecurity pilotproof of concept

    September 3, 2026 · 8 min read · By Fensivo Team

    How to run a human risk platform pilot

    Leer en español

    A good pilot of a human risk management (HRM) platform is not measured by how many employees finished a training module, nor by how much the click rate fell on the first send. It is measured by one thing: whether people's behavior changed when they were put back under the pressure of a simulated attack. And the first step is not choosing a vendor, it is writing down, before you start, the concrete result that would justify the purchase. Everything else, the duration, the population, the metrics, follows from that definition.

    The one-sentence conclusion: what a good pilot must prove

    A pilot exists to answer a question, not to fill an activity dashboard. The right question is whether the platform gets a person who fell for a trick to stop falling for the next one, not whether the team clicked on fewer emails during the trial month. That distinction looks subtle and changes everything, because almost every tool shows an initial drop in click rate and almost none proves the drop holds.

    The reason this matters is that the risk being tested is human, not technical. Cisco's 90-5-5 framework, which estimates that close to 90 percent of breaches involve a human factor, puts the problem where it really sits: in how a person decides under pressure, not in whether they watched a slide. A pilot that measures knowledge (did they pass the course?) answers the wrong question. One that measures behavior (did they fall again when we tested them once more?) answers the one that decides the purchase.

    Human risk is managed automatically.

    Turn human risk into your first line of defense.

    Book a demo

    Free demo · 30 minutes · No commitment

    What to measure in the pilot (and why click rate is not enough)

    The click rate of a single send measures a moment, not learning. A person may skip an email because they were distracted, because it did not apply to their role, or because they happened to be suspicious that day, and fall for the next one. That is why the number that truly matters in a pilot is not the one from the first shot.

    There is peer-reviewed evidence that backs this directly: finishing a training does not on its own predict a reduction in real failures (Ho et al., IEEE Symposium on Security and Privacy 2025; Lain et al., IEEE S&P 2022). What proves change is testing the behavior again. That second test, with a different template but the same type and difficulty, is called a retest, and it is the central metric of a serious pilot.

    Email is still the ground where you test, because more than 90 percent of successful cyber-attacks start with a phishing email (CISA): if the pilot simulates through the channel attacks actually use, it measures what will happen in production.

    Alongside the retest result, it helps to watch the reporting rate (did people start flagging suspicious emails instead of only avoiding them?) and the evolution of each person's risk score across the weeks. On exactly what to watch and why click rate alone misleads, we go deeper in report rate, click rate and retest. The pilot rule is simple: if a metric cannot tell "remembered that email" apart from "changed how they react," it is useless for deciding.

    How long it should last to see behavior change, not the initial scare

    The first simulation almost always lowers the click rate, and that drop is the initial scare, not learning. People go on alert because they know they are being tested. A pilot that ends there buys the illusion of change, not the change.

    To see real behavior, the pilot has to last long enough to close at least one full cycle: a simulation that reveals who falls, immediate remediation for whoever failed, and the retest three weeks later to confirm the lesson held. That puts the realistic floor of a pilot around six to eight weeks, not two. Less than that measures the reflex of the first scare; much more than that starts to look like the rollout itself. The question that sets the duration is direct: is there enough time for the same person to fail, be remediated and be tested again? If the answer is no, the pilot is too short to conclude anything.

    Who to include: why the pilot needs enough mass

    A human risk pilot needs statistical mass for the per-person score to mean something. Below roughly 25 people, what you are measuring is teams, not individuals, and the value of a platform like this lies precisely in telling the person apart, not in averaging the group. A pilot of eight people can feel nimble and say almost nothing.

    The population should not be a comfortable corner of the company either. A good pilot includes a representative sample: a mix of roles and departments, and especially the high-exposure roles, such as finance, leadership and the help desk, which are the ones an attacker goes after first. If the pilot only covers the technology team, usually the hardest to fool, the result looks better than it is and does not represent the organization that will buy.

    Success and decision criteria, defined before you start

    The most expensive mistake in a pilot is deciding what counted as success after seeing the numbers. When the criterion is set at the end, there is always a metric that looks good and serves to justify what you already wanted to do. That is why the criterion is written first. In order:

  1. Write the question the pilot must answer, in one sentence (for example: does the platform get those who fall to stop falling on the retest?).
  2. Set the baseline with the first simulation: who falls, in which deception category and with what starting risk score.
  3. Define the success metric before looking at results: the retest result and the reporting rate, not the click rate of the first send.
  4. Bound the duration and the population: the full cycle (simulation, remediation, retest) over a representative sample of at least 25 people with the high-exposure roles included.
  5. Write the decision criterion (go or no-go) next to the success one: what concrete result enables the rollout and what result stops it.
  6. With those five points written and signed before you start, the pilot stops being a sales demo and becomes a test you can learn from even if it goes badly. If you want to press the vendor with the right questions before signing, we gathered the ones that matter most in ten questions to ask a human risk vendor.

    From pilot to rollout: what each result enables

    A well-designed pilot does not have a single possible outcome, it has several, and each enables a different decision. If behavior improved on the retest, meaning those who fell stopped falling when tested again, the pilot validated what mattered and the rollout is justified with evidence, not faith. If the first send's click rate dropped but the retest failed again, the pilot revealed something valuable: that tool measures activity, not change, and it is worth looking elsewhere. And if the exposed-credential monitoring found compromised accounts on day one, it already delivered a concrete return before a single simulation was measured.

    That last point connects to why the midmarket increasingly adopts through a pilot: it is a fast-growing category, with small and midsize companies as the fastest-growing segment according to Mordor Intelligence, and a pilot is how a midsize company validates the spend before committing. The pilot is not a formality before the purchase, it is the best tool a buyer has to avoid a wrong call.

    At Fensivo we design the pilot on the capabilities that already run: implementation is via OAuth with Google Workspace or Microsoft 365, so the system is live in a day and the first executive report on credential exposure arrives within 48 hours, without waiting to accumulate behavior data. From there the cycle tests with email phishing simulations, remediates on whoever fails and validates with the retest three weeks later, within the 25-to-500-employee focus where the per-person score is valid. You can see how this plays out in concrete situations in our use cases.

    The question we leave open is uncomfortable on purpose: if you had to decide today on the platform you evaluated last quarter, could you say whether your people really changed, or only how many finished the course?

    Sources and references

    Cisco, The 90-5-5 Concept: Your Key to Solving Human Risk in Cybersecurity, 2025: https://blogs.cisco.com/security/the-90-5-5-concept-your-key-to-solving-human-risk-in-cybersecurity CISA, 4 Things You Can Do To Keep Yourself Cyber Safe: https://www.cisa.gov/news-events/news/4-things-you-can-do-keep-yourself-cyber-safe Ho et al., IEEE Symposium on Security and Privacy 2025: https://ieeexplore.ieee.org/document/11023357 Lain et al., IEEE Symposium on Security and Privacy 2022: https://ieeexplore.ieee.org/document/9833766 Mordor Intelligence, Security Awareness Training Market: https://www.mordorintelligence.com/industry-reports/security-awareness-training-market

    Human risk is managed automatically.

    Turn human risk into your first line of defense.

    Book a demo

    Free demo · 30 minutes · No commitment