Teachomatic what to automate, and what not to

False Positives, and Who They Hit

A 1% error rate sounds like nothing. Here is what it looks like from inside an institution.

When Turnitin launched its AI detector in April 2023, the company stated a 1% false positive rate. Vanderbilt University did the arithmetic against its own volume: it had submitted 75,000 papers in 2022, so at that rate around 750 student papers could have been incorrectly labelled as containing AI writing. In August 2023 the university disabled the tool, publishing its reasoning — the false positive problem, the reported cases of wrongly accused students elsewhere, the finding that detectors flag non-native English writers more often, and the fact that Turnitin had not explained how its determination is made beyond saying it looks for patterns it does not define.

For comparison with workplace systems that formalise time or activity data, Monitask’s page on employee PC activity tracking shows how the same type of measurement is handled outside education.

That is the whole argument in one paragraph, and it came from a university rather than from anyone selling anything.

The errors are not spread evenly

This is the part that turns a technical limitation into an equity problem.

For an external perspective on academic integrity, technology, and student rights, see Patterns.

Second-language writers. The core finding is Liang et al. in Patterns, 2023: seven detectors, 91 supervised TOEFL essays, a mean false-positive rate of 61.3%, while essays by US eighth-graders were classified almost perfectly. The mechanism is that these tools reward elaborate prose, and writing in a second language is not elaborate.

First-generation students. Later academic work examining detector bias lists first-generation students alongside non-native speakers as disproportionately affected — the same mechanism, a different route to plainer prose.

Neurodivergent students. Weaker evidence, and worth labelling as such. What exists is largely institutional observation rather than controlled study: the University of Nebraska-Lincoln has been reported as finding more false positives among neurodivergent students. The proposed explanation is plausible — writing that is unusually regular or formulaic reads as machine-like — but nobody should quote a percentage here, because there is not a reliable one.

Anyone taught to write in a formulaic structure. Which is to say, anyone drilled for an exam. The five-paragraph essay is a low-perplexity artefact by design.

Notice the pattern. Every group on that list is one that a school is otherwise trying to support.

What institutions have done

Vanderbilt is the canonical case, and it was not alone or last. The University of British Columbia declined to enable Turnitin's detection within days of its launch in April 2023. UC Berkeley declined partly on privacy grounds — the question of feeding student work to a third party. The University of Waterloo cited peer-reviewed studies and an internal test in which human-written text was flagged as fully AI-generated. Curtin University disabled it from 1 January 2026, framing the decision as education rather than surveillance.

Trackers of these decisions now count somewhere above fifty institutions across several countries. Take the totals with care — most such trackers are maintained by companies with an interest in the conclusion — but the individual cases above are documented by the universities themselves.

Two things follow for a school still using one. First, the position that detection is standard practice is no longer accurate. Second, the reasons those institutions gave are available in writing and are useful in a meeting.

Why "1%" was never the right number to argue about

Even taking a vendor's figure at face value, two things make it misleading.

The rate is measured on a population, and applied to a person. Vanderbilt's arithmetic makes this concrete: 1% is small, 750 students is not.

And the rate that matters is not the one advertised. What you need to know is: given that this paper was flagged, how likely is it that this student did nothing wrong? That depends on how common undeclared AI use actually is in your cohort — and the rarer it is, the higher the proportion of flags that are innocent. With a rare behaviour, a good detector still produces a pile of accusations you cannot stand behind individually.

What this means for a teacher who is not making policy

You may work somewhere that runs a detector and does not intend to stop. Three things are within your control.

Never let a score be the reason on its own. What you owe a student before raising it does not change because the school bought a tool.

Tell your class you will not act on a score alone, and mean it. Many students are quietly frightened of these tools; some are running their own work through detectors before submitting to pre-empt an accusation. Saying this out loud costs nothing and changes the room.

Record the ones that dissolve. Schools almost never learn how often their tool is wrong, because a flag that turns out to be nothing leaves no trace. If you keep a private count over a term, you will have the only local evidence anyone has.

And the design answer

The durable version of this is not better detection. It is setting work where the question does not arise — anchored to your room, to the student, or to their reasoning.

Detection asks a machine to establish authorship after the fact, on evidence it cannot explain. Assessment design makes authorship visible while the work is happening. Only one of those has ever worked.

The short version