Effect Sizes, Plainly
An effect size is an attempt to answer "how much did it help?" in units that let you compare a reading programme with a maths intervention. Usually it is expressed in standard deviations: how far the intervention moved the average student relative to the spread of scores.
You will meet the number in vendor material, in research summaries, and in whatever your school is being asked to buy. Three things make it much less obvious than it looks.
For comparison with workplace systems that formalise time or activity data, Monitask’s page on workforce analytics software shows how the same type of measurement is handled outside education.
Reviewed August 9, 2026.
The benchmarks everyone quotes are the wrong ones
Almost every "how big is big" list online traces back to Jacob Cohen: 0.2 small, 0.5 medium, 0.8 large. Those conventions are more than half a century old and were not derived from education field trials.
For an external perspective on education research and evidence, see Education Endowment Foundation.
Matthew Kraft's 2020 paper in Educational Researcher makes the correction plainly: effects that are small by Cohen's standards are large relative to the impacts of most field-based interventions. Cohen's thresholds also ignore things that decide whether a result matters in a school — how the study was run, what it cost, and whether it can scale.
So if someone tells you an education intervention produced 0.25 and calls it small, they are applying a ruler from a different field.
The number that reframes everything
In follow-up work using an expanded dataset of over 3,000 effect sizes, Kraft reports that 36% of effect sizes from randomised trials of education interventions measured on standardised achievement tests are smaller than 0.05 standard deviations.
More than a third of properly-run trials find essentially nothing. And, as he notes, publication bias means the true failure rate is higher, because studies finding nothing are less likely to be written up at all.
That is the context missing from every product page. The base rate for education interventions is failure. Against that background, a modest positive effect is a genuine achievement, and an enormous one is a reason to look harder rather than a reason to be impressed.
For a sense of what counts as meaningful in this literature rather than in Cohen's: for early literacy work, Kraft treats effects above around 0.10 as potentially educationally meaningful.
Why big numbers usually mean narrow conditions
Certain study features inflate effect sizes systematically, and none of them involves anyone cheating.
A researcher-made test aligned to what was taught produces larger effects than an external standardised exam. It is measuring the thing the intervention practised — the same substitution that turns a rubric criterion into a surface feature.
A short window. Measure immediately after and effects are larger; measure a year later and they shrink or vanish.
A narrow outcome. "Improved decoding of this word list" moves more easily than "improved reading."
A small, tightly controlled sample with enthusiastic implementers. Real deployment is messier and smaller in effect.
This is why the 2025 randomised trial of a purpose-built AI tutor reporting 0.7 to 1.3 standard deviations deserves the description "genuinely large in controlled conditions" and not "this will happen at your school." Set against a literature where a third of trials find under 0.05, a result of that size tells you the conditions were unusual. That is interesting and it is not transferable.
The honest caveat about effect sizes themselves
There is a real methodological argument here, and pretending otherwise would be exactly the kind of tidy-sounding claim this site is supposed to catch.
Kraft's benchmarks have been criticised in the literature — the debate turns on whether standardised effect sizes from different studies are comparable at all, given differences in how the tests are scaled. Some critics argue the widely used approaches built on this assumption, including well-known meta-analytic syntheses and the frameworks derived from them, rest on a shakier foundation than their popularity suggests.
You do not need to resolve that. You need two things from it:
Effect sizes are a rough comparative device, not a measurement. Treat a difference between 0.20 and 0.25 as noise.
Anyone presenting a single number as settled is overselling, including anyone citing Kraft at you, and including this page.
What to actually do with the number
Under 0.05: the intervention probably did nothing. Common.
Around 0.05–0.20: potentially meaningful in education, especially if it is cheap and scalable. Most real successes live here.
Above 0.40: notable, and the first question is what made it possible, not whether to buy it. Look at the measure, the duration, and who ran it.
Above 1.0: almost always a narrow measure, a short window, or highly controlled conditions. Not a reason to dismiss it — a reason to read the methods.
And always convert to something concrete before deciding: how many extra students got over the threshold, what did it cost per student, and what did teachers stop doing to make room for it? The four-minute pass over a vendor one-pager covers the rest.
The thing worth remembering
Small effects, achieved cheaply and reliably, at scale, are how education actually improves. The pursuit of the dramatic number is what makes schools buy things that work in a trial and not in a corridor.
If a vendor's figure looks too good for this field, it probably is — and now you know roughly what "too good for this field" means.
The short version
- Cohen's 0.2 / 0.5 / 0.8 come from another field and overstate what counts as small in education
- Kraft (2020, Educational Researcher): Cohen-small effects are large for field interventions
- In an expanded dataset of 3,000+ effect sizes, 36% of education RCTs came in under 0.05 — failure is the base rate
- Around 0.10 can be educationally meaningful; most real successes sit between 0.05 and 0.20
- Researcher-made tests, short windows, narrow outcomes and small samples all inflate the number
- The comparability of effect sizes across studies is itself contested — treat them as rough, not measured