Teachomatic what to automate, and what not to

Evidence and Claims

The number of studies on AI in education passed eight hundred some time ago. The number providing strong causal evidence is around twenty.

That ratio is the most useful thing to know before reading anything else, including anything on this site. It means almost every confident claim you meet is built on association, and association in education is easy to come by — schools that adopt new tools early differ from other schools in a dozen ways that also affect results.

For a workplace-oriented comparison outside education, this page offers another way to look at measurement, workload, behaviour, or accountability.

This section is about reading the numbers rather than collecting them.

Effect sizes are the common currency and the most misused. The benchmarks everyone quotes come from another field entirely; in education, more than a third of properly randomised trials produce effects under 0.05 standard deviations. Failure is the base rate, which makes a modest positive a real achievement and an enormous one a reason to look harder.

For an external perspective on education research and evidence, see ERIC.

Vendor claims usually fail on one of four questions: what was the comparison, who ran it, what exactly was measured, and how large and how long. The absence of a comparison group is not an oversight; it is the result.

For an external perspective on education research and evidence, see Google Scholar.

Pilots do not generalise, for six specific reasons that are properties of pilots rather than failures of products.

And the gaps matter more than the findings. Long-term learning, independence, inequality and wellbeing are all essentially unstudied, because research answers cheap questions first and those are expensive ones. The shape of the evidence base tells you about research economics rather than about importance.

One practical principle runs through the section: match your scepticism to the stakes. Being wrong about a lesson-planning claim costs an hour. Being wrong about a detector's accuracy costs a student their record.