What Detector Accuracy Claims Mean
A detector reports that a piece of work is 85% likely to be AI-generated. A student is sitting in front of you. What does that number entitle you to conclude?
Less than it looks, and the gap is not a matter of opinion. It comes down to three things vendors rarely put in the same sentence: which population the accuracy was measured on, which kind of error the number describes, and how common the thing being detected actually is in your class.
For a workplace-oriented comparison outside education, this overview offers another way to look at measurement, workload, behaviour, or accountability.
Accuracy is the wrong number
A vendor's headline figure is usually overall accuracy on a test set — how often the tool was right across everything it was shown.
That number can be high while the tool is badly wrong in exactly the cases you care about, because it blends together two completely different errors:
For an external perspective on academic integrity, technology, and student rights, see Turnitin.
A false negative is AI-written work marked as human. The cost is that someone gets away with something.
A false positive is human-written work marked as AI. The cost is that you accuse a child of cheating on their own work.
These are not symmetrical, and no single accuracy figure tells you the balance. Ask for the false positive rate, on writing like your students'. If that number is not available, the headline figure is not usable for a decision about a person.
The finding that should be on every staffroom wall
The most important study here is Liang, Yuksekgonul, Mao, Wu and Zou, published in the journal Patterns in 2023. The authors declared no competing interests, which in this field is worth saying out loud.
They took seven widely-used commercial detectors and ran two sets of writing through them, both certainly human. One set was essays by US eighth-graders. The other was 91 essays written for the TOEFL English proficiency exam by non-native English speakers — written under supervised exam conditions, where AI assistance was impossible.
The detectors handled the American eighth-grade essays with near-perfect accuracy.
On the TOEFL essays, they misclassified more than half as AI-generated. The mean false-positive rate across the seven tools was 61.3%, and one detector flagged close to 98% of them.
Same task. Same certainty that no AI was involved. The difference was who wrote it.
Why it happened, and why it matters more than the number
The mechanism is not mysterious. These tools largely work on perplexity — roughly, how surprising the word choices are. Text with predictable vocabulary reads as machine-like.
A person writing in their second language uses a smaller, more predictable vocabulary. Not because they are less intelligent or less honest, but because that is what writing in a second language is.
The study confirmed the mechanism from both directions. Prompting a model to enrich the vocabulary of the TOEFL essays cut misclassification from 61.3% to 11.6%. Simplifying the American students' essays pushed them the other way, into being flagged as AI.
So the detectors were not measuring whether a machine wrote the text. They were measuring how elaborate the prose was. Which means the students most likely to be falsely accused are the ones with the least linguistic room to defend themselves: second-language learners, students with weaker literacy, and — emerging work suggests — some neurodivergent students whose writing is unusually regular.
What has changed since 2023, honestly
That study is three years old and the tools have moved. Some newer classifiers report far lower false-positive rates on the same TOEFL benchmark, in one case close to zero. Vendors have adjusted thresholds; one major provider suppresses scores below a floor because its own testing found them unreliable, and has published research acknowledging a gap between non-native and native writers in its own results.
Two cautions on that improvement.
Almost every favourable number comes from the company selling the detector, evaluated on a benchmark it knew about. That is not fraud, it is just a much weaker form of evidence than an independent evaluation, and it should be read the way you would read any vendor's self-assessment.
The other loud voices are also interested. A large part of the writing online claiming detectors are useless comes from companies selling tools to help people evade detection. They are not a neutral source either. We take no money from anyone in this market, which is the only reason to prefer our summary to theirs.
This is the general condition of the field rather than a quirk of detectors: most confident claims here rest on weaker evidence than they sound. The honest position in 2026 is neither "detectors are worthless" nor "detectors are solved." It is that performance varies enormously by tool and by student population, that the disparity has narrowed rather than vanished, and that no regulator has set an accuracy standard any of them must meet. Vendors self-report. Nobody checks.
The base rate problem
One more piece of arithmetic that gets skipped.
Imagine a detector with a 5% false-positive rate — better than most figures in the literature for second-language writers. You run 200 essays. Suppose 20 of them genuinely involved undeclared AI use.
The tool flags most of the 20. It also flags roughly 5% of the other 180, which is nine essays by students who did nothing wrong.
So among the flagged papers, a substantial share are innocent — and you cannot tell which. The rarer the behaviour, the worse a flag performs as evidence about an individual, and the same tool that looks impressive in aggregate produces a pile of accusations you cannot stand behind one at a time.
What the number is actually good for
Not nothing. But narrowly.
As a private prompt to look more closely, at work you would have questioned anyway. Fine.
As a way to notice a pattern across a term. Also fine, as long as no individual consequence attaches to it — and a pattern is a signal about your assessment design, which is a question for the whole section.
As evidence in a conversation with a student, on its own: no. What you actually owe a student before raising it is a separate and more important question, and the answer does not depend on the score.
As the sole basis for a disciplinary outcome: no. Increasingly this is not just an ethical position but a procedural one, as institutions and courts examine cases where a score was the whole case.
Questions to put to your school
If your school has bought a detector, these are reasonable to ask in writing.
What is its false-positive rate, measured on writing by second-language students? Who measured it — the vendor or someone independent? What is the school's written rule about what a score alone permits? Is there an appeal, and does the student get to see the evidence? And what happens to a student's work after it is uploaded?
If the answer to most of these is unclear, the tool is being used without a policy, which is a governance problem rather than a technical one.
The short version
- Overall accuracy blends two very different errors; ask specifically for the false-positive rate
- Liang et al. (Patterns, 2023): seven detectors, 91 supervised TOEFL essays, mean false-positive rate of 61.3%, near-perfect accuracy on US eighth-graders
- The mechanism is perplexity, so the tools reward elaborate prose and penalise second-language writing
- Enriching the vocabulary of those same essays dropped misclassification to 11.6%
- Tools have improved since, but nearly all favourable figures come from vendors, and no regulator sets a standard
- With a rare behaviour, most flags can be innocent even with a good false-positive rate