Teachomatic what to automate, and what not to

Why Pilots Don't Generalise

A pilot runs in three classrooms. It goes well. The tool is bought for the whole school, and a year later nobody can say whether it helped.

This is such a common sequence that it deserves an explanation better than "people lost interest." There are six specific mechanisms, and most of them are properties of pilots rather than failures of anyone involved.

For a workplace comparison of the underlying concept, Monitask’s discussion of self-reporting bias shows how the same behaviour is framed outside education.

The six mechanisms

Volunteers. The teachers in a pilot chose to be there. They are more motivated, more tolerant of friction, and more likely to make something work through sheer commitment. Rollout replaces them with everybody, including the exhausted, the sceptical, and the person covering two subjects. The volunteer effect is large and has nothing to do with the product.

Attention. During a pilot, someone is watching. There are check-ins, a contact at the vendor, a senior leader who cares. Attention improves outcomes independently of what is being attended to, and it is the first thing withdrawn at scale.

For an external perspective on education research and evidence, see Brookings education research.

Support that does not scale. Setup done by the person who understood it, accounts configured by hand, problems fixed within a day. At three classrooms this is invisible. At forty it is a job nobody has.

Selection of classes. Pilots run where conditions are good — a cooperative group, a subject that fits, a room with working devices. Not dishonesty; just the natural instinct to give it a fair chance. The result is a measurement of the best case.

Novelty. New things produce engagement for a few weeks in both directions: students are interested, teachers are energised. Any measure taken inside that window includes the novelty, and novelty does not renew.

The short window. Six weeks is enough to see enthusiasm and not enough to see whether it survives a term with exams, illness, cover, and a broken projector.

None of these requires anyone to be misleading. They are why a pilot's result is a claim about the pilot.

Reading someone else's pilot

If a vendor's evidence is a pilot — and it usually is — three questions get you most of the way. The wider four-minute pass covers the rest.

Who was in it, and how were they chosen? Volunteers, or everybody in a department?

What support did they get that we would not? If the answer is a dedicated contact and hands-on setup, subtract accordingly.

How long, and when was it measured? Immediately after a six-week pilot is the most flattering possible moment.

And keep the base rate in view: more than a third of properly run trials in education find almost nothing. A pilot that reports a clear positive is, statistically, more likely to be a pilot artefact than a discovery.

Running one that tells you something

Pilots are still worth doing. They are just usually designed to answer the wrong question — did it work? — when the useful question is what would it take?

Include somebody who did not volunteer. One reluctant participant tells you more than three enthusiasts. Ask them, not to be difficult, but because their experience is the one that predicts rollout.

Write down what success would look like before you start. Specifically, and including what would count as failure. Otherwise you will find a positive, because there is always a positive available afterwards.

Log the friction, not just the outcome. Every time something did not work, every workaround, every minute of setup. This is the data that predicts rollout and it is the data nobody collects.

Run it long enough for novelty to wear off. A term, minimum. Measure at the end, not in week three.

Include a class where conditions are bad. Old devices, a difficult group, a subject that fits awkwardly. If it survives there, you have learned something.

Ask what teachers stopped doing. Time came from somewhere. If it came out of marking or planning, that is a cost, and it belongs in the evaluation.

The question that decides rollout

Not did it work but what would have to be true for this to work everywhere here?

Devices for every student. Twenty minutes of training that people actually attend. Someone whose job includes fixing it. A subject shape it suits. Then check honestly whether those things exist outside the pilot.

Usually one of them does not, and that is the answer. Better to know it before the purchase order than after.

And your own classroom

The same logic applies to a sample of one.

If you try something and it works, the useful test is two weeks measured by hand, across a normal fortnight rather than an enthusiastic one — including the week you are behind and the day nothing loads.

Anything that only works when you are on top of things is not a tool you have. It is a tool you borrow when things are already going well.

The short version