What the Studies Do and Don't Show
The number of studies on AI in education passed eight hundred some time ago. The number that provide strong causal evidence — where you can say the AI caused the outcome rather than accompanied it — is around twenty.
That ratio is the single most useful thing to know before reading anything else in this field, including anything on this site. It means almost every confident claim you encounter is built on association, and association in education is very easy to come by, because the schools that adopt new tools early differ from other schools in a dozen ways that also affect results.
For a workplace-oriented comparison outside education, this overview offers another way to look at measurement, workload, behaviour, or accountability.
Reviewed August 9, 2026.
The one finding that holds up
AI saves teachers time on preparation and first-draft feedback.
For an external perspective on education research and evidence, see Science.
This is where the evidence is most consistent, and it is worth noticing why: it is the easiest thing to measure, the effect appears quickly, and the outcome — hours — does not require a contested definition. One study put the saving at roughly half an hour a week without a drop in quality. Survey work in 2026 puts the proportion of teachers reporting improved practice around two-thirds, though surveys measure perception rather than effect.
Half an hour a week is not transformative and the honest framing matters here. In an overloaded job it is real. It is also considerably less than the marketing implies.
And there is a documented catch. Time saved does not reliably become time free. In several studies teachers spent the recovered time giving more or deeper feedback. That is arguably a better outcome than an early finish, but it means "this will reduce your workload" is not what the evidence says. It says the shape of the work changes.
The one finding that is promising and unsettled
Purpose-built tutoring systems can produce large learning gains.
The strongest single result here is a 2025 randomised controlled trial reported in Scientific Reports, in which a purpose-built AI tutor produced gains in the range of 0.7 to 1.3 standard deviations over active learning, with students covering more in less time. In education research those are very large numbers indeed — more than a third of trials in this field find under 0.05.
Three reasons not to generalise from it.
It was a purpose-built system, not a chatbot. The design of the pedagogy did the work, not the model underneath. A general assistant used the same way is a different intervention.
Design determines direction. Across studies, systems that ask questions, steer, and give step-by-step guidance perform differently from systems that simply supply answers. AI-as-tutor and AI-as-cheat-sheet are opposite interventions that look identical from a distance.
One trial is one trial. A striking result in a controlled setting is a reason to pay attention, not a reason to buy anything. The full reading of the tutoring evidence sets out what it does and does not support.
What we genuinely do not know
This list matters as much as the findings, and it is the part vendors never print.
Long-term effects on deep learning. Almost nothing. Studies run for weeks; the questions people actually care about run for years.
Effects on metacognition and independence. Whether routine assistance erodes a student's ability to work through difficulty alone is the most commonly raised worry in the debate and among the least studied.
Effects on wellbeing and social development. Barely touched.
Effects on inequality. Genuinely open, and it could go either way. Adaptive support might help students with the least help at home, or better-resourced schools might extract more from the same tools. Both stories are plausible and neither is settled.
Notice that this list is roughly the set of questions any thoughtful parent asks first. That is not a coincidence — the easy-to-measure questions get answered first, and they are not the important ones.
How to read a claim in this field
Four questions, applicable to anything.
Causal or correlational? "Schools using X see higher scores" is almost always the second, and almost always presented as the first.
Compared with what? Against nothing? Against business as usual? Against active learning by a good teacher? The comparison decides the number, and the flattering comparison is the common one.
Who paid, and who ran it? A vendor-run pilot with vendor-defined success measures is marketing with a methods section. It can still be true; it is just much weaker evidence. Applying these four questions to a specific one-pager takes about four minutes.
What was actually deployed? "AI improved outcomes" tells you nothing. Which system, doing what, in whose hands, for how long?
Reading a specific vendor claim is a longer exercise, and the four questions above get you most of the way.
What this means for your Monday
Not paralysis. A specific division of confidence.
Act on the teacher-side findings. Prep, differentiation, first drafts of feedback — supported, low-risk, and you can check the output yourself. Where the time actually goes is the practical version.
Treat student-side claims as hypotheses. Try things, watch what happens in your own room, and do not let a vendor's chart substitute for that.
Be sceptical in proportion to the stakes. A weak claim about lesson planning costs you an hour. A weak claim about a detector's accuracy costs a student their record — which is why that evidence gets examined separately and much harder.
The short version
- Over 800 studies, roughly 20 with strong causal evidence — assume association unless told otherwise
- Best-supported finding: teachers save time on prep and first-draft feedback, around half an hour a week in one study
- Saved time often becomes deeper feedback rather than fewer hours; the work changes shape
- Large tutoring gains exist in controlled trials of purpose-built systems, not of general chatbots
- Long-term learning, metacognition, wellbeing and inequality effects are essentially unstudied
- Ask: causal or correlational, compared with what, who paid, and what was actually deployed