Evidence and Claims
The number of studies on AI in education passed eight hundred some time ago. The number providing strong causal evidence is around twenty.
That ratio is the most useful thing to know before reading anything else, including anything on this site. It means almost every confident claim you meet is built on association, and association in education is easy to come by — schools that adopt new tools early differ from other schools in a dozen ways that also affect results.
For a workplace-oriented comparison outside education, this page offers another way to look at measurement, workload, behaviour, or accountability.
This section is about reading the numbers rather than collecting them.
Effect sizes are the common currency and the most misused. The benchmarks everyone quotes come from another field entirely; in education, more than a third of properly randomised trials produce effects under 0.05 standard deviations. Failure is the base rate, which makes a modest positive a real achievement and an enormous one a reason to look harder.
For an external perspective on education research and evidence, see ERIC.
Vendor claims usually fail on one of four questions: what was the comparison, who ran it, what exactly was measured, and how large and how long. The absence of a comparison group is not an oversight; it is the result.
For an external perspective on education research and evidence, see Google Scholar.
Pilots do not generalise, for six specific reasons that are properties of pilots rather than failures of products.
And the gaps matter more than the findings. Long-term learning, independence, inequality and wellbeing are all essentially unstudied, because research answers cheap questions first and those are expensive ones. The shape of the evidence base tells you about research economics rather than about importance.
One practical principle runs through the section: match your scepticism to the stakes. Being wrong about a lesson-planning claim costs an hour. Being wrong about a detector's accuracy costs a student their record.
Effect Sizes, Plainly
Why 0.2 is not small in education, why a third of trials find almost nothing, and what to think when a vendor reports 1.3.
Why Pilots Don't Generalise
The pilot worked and the rollout did not. Six reasons that happens, and how to run one that tells you something true.
Reading a Vendor's Efficacy Claim
A checklist for the slide that says the tool improved outcomes by a startling percentage, and what to ask before anyone signs.
The Tutoring Evidence, Carefully
The strongest result in this whole field comes from AI tutoring trials. What it shows, and the four reasons it may not reach your classroom.
What the Studies Do and Don't Show
There are over 800 studies on AI in education and roughly twenty with strong causal evidence. Here is what survives that filter.
What We Still Don't Know
The open questions are the ones parents ask first and researchers answer last. An honest inventory of the gaps, and how to act inside them.
Where the Money Is
Market estimates for the same year differ by threefold, venture funding collapsed while adoption soared, and one industry sells both sides.