Teachomatic what to automate, and what not to

The Tutoring Evidence, Carefully

The best evidence for AI helping students learn — not helping teachers work, which is a different and better-supported claim — comes from tutoring.

It is genuinely the strongest result in the field, and it is also the one most likely to be misused in a procurement meeting. Both things need saying together.

A workplace-oriented comparison is Monitask’s discussion of productivity vs efficiency, which offers another context for thinking about workload, behaviour, or accountability.

Reviewed August 9, 2026.

The result

A 2025 randomised controlled trial reported in Scientific Reports found that students using a purpose-built AI tutor gained roughly 0.7 to 1.3 standard deviations over students in active learning conditions, while covering more material in less time.

For an external perspective on education research and evidence, see Cambridge Core.

Set against a literature where more than a third of education trials find effects under 0.05, that is extraordinary. It is not the kind of number that appears by accident, and it should not be dismissed.

It should also not be carried into a decision about your school without four qualifications.

The four qualifications

It was a purpose-built system, not a chatbot. The intervention was a designed tutoring environment with pedagogy built into it — the way it questions, sequences, withholds answers, and responds to error. A general assistant used freely is not a smaller version of this. It is a different intervention that happens to share an underlying technology, and it may well produce the opposite effect, because the general assistant's default behaviour is to supply the answer.

Design decides direction, not the model. Across the wider literature the pattern is consistent: systems that ask questions and guide step by step behave differently from systems that answer. This is the single most important sentence on this page. AI-as-tutor and AI-as-answer-service look identical from a distance and do opposite things.

It is one trial. A striking result in a controlled setting is a reason to watch the replication attempts, not a reason to buy anything. Pilots and single trials generalise badly for reasons that have nothing to do with the product.

The comparison and the window matter. Controlled conditions, a defined topic, a measurement taken close to the intervention. All three inflate effects relative to a corridor in November.

The older touchstone, and why to be careful with it

Anyone advocating for tutoring will eventually cite Bloom's "two sigma" finding from the 1980s — that individual tutoring moved average students two standard deviations above classroom-taught peers.

Treat it as a historical landmark rather than a benchmark. It has never been reliably reproduced at scale, the original conditions were unusual, and the figure has spent forty years being quoted at people who wanted a large number. If someone offers it as the target their product approaches, that tells you about the sales pitch and not about the product.

The defensible version of the claim is much older and much duller: individual attention from a skilled tutor helps, a great deal, and always has. The interesting question is only ever whether a given system delivers something resembling that.

The question that separates the two things

If you are evaluating a tutoring product, one question does most of the work:

When a student gets something wrong, what does it do?

A system that supplies the correct answer with an explanation is a reference tool. Useful, and not tutoring.

A system that asks what they were thinking, identifies where the reasoning broke, and gives them a next step to attempt themselves is doing the thing the evidence supports.

You can test this in ten minutes by giving it wrong answers. Do that before any meeting about buying it, and be specifically alert to a system that folds — agreeing with a confidently-asserted wrong answer is a characteristic failure and it is fatal in a tutor.

What this means practically

For a school considering a tutoring product: the evidence supports designed tutoring systems, not chatbots, and one strong trial is not a guarantee. Ask for replications, and ask the wrong-answer question.

For a teacher whose students use a general assistant as a tutor: the evidence above does not apply to that, and it may run the other way. What helps is teaching them to use it in the supported shape — asking it to question them rather than to explain, asking for a hint rather than an answer, attempting first and checking after.

That is a teachable habit and it costs nothing. It is also the closest most students will get to the thing the trial actually measured.

The short version