Reading a Vendor's Efficacy Claim
Somebody forwards a one-pager. A tool improved outcomes by 34%. There is a chart, a school district, a quotation from a head of department.
You are not going to read the underlying study, because there usually is not one and you have lessons. So here is what to do in the four minutes you actually have.
For a workplace-oriented comparison outside education, this guide offers another way to look at measurement, workload, behaviour, or accountability.
None of this assumes bad faith. Most vendor claims are produced by people who believe them, though the shape of the market explains which claims reach you. The problem is structural: every incentive in the process pushes toward the flattering version, and nobody in the chain is paid to find the unflattering one.
The four-minute pass
1. Find the comparison. Improved by 34% — compared with what? Against the same students last year, against a different school, against doing nothing at all? The comparison is where most of the effect lives, and the vaguer the sentence, the weaker it usually is. If there is no comparison group, there is no claim, only a before-and-after in a year when other things also changed.
For an external perspective on education research and evidence, see JSTOR.
2. Find who ran it. Vendor-run, vendor-funded, or independent? A vendor-run pilot with vendor-chosen measures is marketing with a methods section. Not worthless — just the weakest tier of evidence, and it should be labelled.
3. Find the outcome that was measured. "Improved outcomes" hides an enormous range. Engagement? Time on task? Teacher satisfaction? Scores on a test the vendor supplied? Scores on an external exam? Only the last two are learning, and only the last is hard to influence.
4. Find the sample. How many students, in how many schools, for how long? The wider evidence base is thin enough that a single pilot carries more weight than it should. A striking result from one enthusiastic school over six weeks tells you about that school and those six weeks.
If a one-pager does not let you answer all four, that is itself the finding. The absence of a comparison group is not an oversight; it is the result.
The five moves to recognise
The unnamed comparison. "Students using X outperformed their peers." Which peers? Selected how? Schools that adopt new tools early differ systematically from those that do not, and that difference alone produces the effect.
The proxy outcome. Engagement and time-on-task are easy to move and easy to measure, and their relationship to learning is weak. A tool that increases time on task has increased time on task.
The pilot with volunteers. Teachers who volunteer for a pilot are more motivated and better supported for the duration. That effect is large and has nothing to do with the product — one of six reasons pilots do not generalise.
The percentage without a base. "34% improvement" on what starting figure, in what units? A 34% rise in a small number is a small number.
The unpublished internal study. "Our research shows." Ask where it is published. The answer is often that it is not, which means it has been reviewed by the people who wanted the result.
Effect sizes, briefly
If a claim reports an effect size, that is a good sign — it means somebody did statistics. Two things to know.
Educational interventions typically produce small effects. Anything above about 0.4 standard deviations is notable, and things far above that usually turn out to involve a narrow measure, a short window, or a comparison against nothing.
Very large effects do occur in tightly controlled conditions. A 2025 randomised trial of a purpose-built tutoring system reported gains of 0.7 to 1.3 standard deviations — genuinely large. But that was a purpose-built system in a controlled setting, and it does not transfer to a general assistant used differently. A real result in one condition is not a claim about your school.
The questions to ask before signing
If your school is deciding, these are reasonable in writing, and the response tells you as much as the answers.
Who conducted the study, and who paid? What was the comparison group and how was it formed? What exactly was measured, and by whom? How many students, over how long? Have there been evaluations showing no effect, and can we see them? What happens to student data — where does it go and who else sees it? And what does the contract say about price after the first year?
A vendor with a solid case answers these easily. A vendor who becomes vague about the comparison group has told you what you needed to know.
The one nobody asks
What would have to be true for this not to work here?
Every product works somewhere. The useful question is what conditions it needs — devices, connectivity, teacher time to learn it, a particular subject shape, a particular age band — and whether you have them.
Ask a vendor where their tool does not work. A good one will tell you, because they would rather not have a failed deployment as a reference site. The same question applies to detectors, where no regulator has set a standard anyone must meet. A vendor who claims it works everywhere for everyone is describing a product that has never been evaluated honestly.
And the claims on this site
The same test applies here.
We take no money from anyone in this market, which removes one incentive and not all of them. Where we cite a figure we say who produced it and when, and where the evidence is thin we say that instead of rounding it up. If you find a number here without a source or a date, that is a mistake and we would like to know.
The short version
- Four minutes: find the comparison, who ran it, what was measured, and the sample size
- No comparison group is not an oversight — it is the result
- Recognise: unnamed comparisons, proxy outcomes, volunteer pilots, percentages without a base, unpublished internal studies
- Above about 0.4 standard deviations is notable; very large effects usually mean a narrow measure or short window
- Ask where the tool does not work; a vendor who says "everywhere" has never been evaluated honestly
- The response to a written list of these questions is itself informative