humanflow

AI detectors for teachers: reading a score before you act on it

A detection score is a statistical estimate, not evidence of anything. Peer-reviewed testing found seven detectors averaged a 61.22% false positive rate on essays written by real people. Treat a high score as a reason to look at drafts and process — never as a finding on its own.

Last reviewed 4 August 2026 · The HumanFlow team

What the number actually is

An AI detector does not know who wrote anything. It has no record of authorship, no browser telemetry and no access to the document's history. It reads the prose and estimates how statistically typical of machine writing it looks — mostly through perplexity, how predictable each next word is, and burstiness, how much sentence length and structure vary.

Human writing tends to be uneven. Model writing tends toward smooth and medium-everything. That difference is the whole mechanism, and it is why the tools struggle with the specific kinds of human writing that happen to be even and conventional.

So a score of 80% does not mean an 80% chance the student used AI. It means the text carries features the model associates with machine writing. Those are not the same claim, and the gap between them is where nearly every unfair accusation lives.

The false-positive research, in plain terms

In 2023, Weixin Liang and colleagues at Stanford published a study in Patterns testing seven detectors against 91 TOEFL essays written by real people — non-native English speakers sitting an English exam. The detectors flagged an average of 61.22% of those essays as AI-generated. Eighty-nine of the 91 were flagged by at least one tool.

The same detectors classified essays by US 8th-grade students almost perfectly. The variable was not whether a machine wrote the text. It was whose English was being read.

Turnitin states its own false positive rate is under 1%, and that figure deserves to be reported accurately: it applies only to documents where at least 20% of the qualifying text is flagged. Below that threshold Turnitin shows an asterisk rather than a percentage, because its own testing found more false positives in that range. Any summary quoting the sub-1% figure without that condition is quoting it wrongly.

Turnitin also funded research testing roughly 2,000 samples from English language learners, reporting false positive rates of 0.014 for ELL writers against 0.013 for native speakers on documents over 300 words. That is a real result and a direct answer to the bias charge. It was also paid for by Turnitin, which does not make it wrong but is worth knowing when the strongest evidence for a tool comes from the company selling it.

The full evidence review is here, including where the studies disagree and why.

Who gets flagged unfairly

The published risk factors are consistent, and they describe students rather than cheating:

  • Non-native English speakers — the single largest documented factor, by a wide margin
  • Students taught rigid structures: five-paragraph essays, formulaic topic sentences, signposted transitions
  • Technical and scientific writers, whose register is deliberately flat and conventional
  • Heavy self-editors, who smooth out exactly the unevenness a detector reads as human
  • Students writing in a second discipline, reaching for safe phrasing because the subject is unfamiliar
  • Anyone using a grammar checker, which pushes prose toward standard, polished English

Notice what these have in common. Every one describes a student whose prose is unusually even before any tool touched it. A detector cannot tell "polished by a careful person" from "polished by a machine" — it only sees polish.

Before you raise it with a student

Work through these in order. Most concerns resolve before the last one.

  1. Check the score clears the tool's own threshold. Turnitin suppresses scores between 0 and 20% for a reason. If your tool marks a result as low confidence, that is the vendor telling you not to rely on it.
  2. Check the document length. Turnitin needs 300 words of prose; most detectors are unreliable on short pieces. A flagged 400-word reflection is much weaker evidence than a flagged 3,000-word essay.
  3. Run it through a second detector. If they disagree sharply, you have learned something about the tools rather than the student.
  4. Ask whether the student is in a documented false-positive group. If English is their second language, the base rate of error is high enough that the score alone should not move you.
  5. Look for process evidence first. Version history in Google Docs or Word, earlier drafts, notes, outlines. This is stronger in both directions than any score, and it is usually available in minutes.
  6. Compare against their previous work. A sudden shift in voice is more informative than a percentage — and its absence is reassuring.
  7. Open with a question, not a verdict. "Talk me through how you approached this" gets you further than presenting a score. A student who wrote the work can almost always discuss their choices; the conversation itself is evidence.

What the conversation costs when it goes wrong

An accusation is not a neutral act. Students who are wrongly accused describe it as one of the worst experiences of their education, and the damage does not undo cleanly when the case is dropped. The base rates above mean that if you act on scores alone across a cohort, you will be wrong about some students — and disproportionately about the ones already at a disadvantage.

A related trap: students who did nothing wrong sometimes confess anyway, to end the discomfort. If your process makes admission the fastest route out of a difficult meeting, it will produce admissions regardless of what happened.

None of this is an argument for ignoring AI use. It is an argument for treating the score as the beginning of an enquiry rather than the end of one.

Designing assessment that needs less detection

The teachers who report the fewest disputes tend to have changed the assignment rather than the enforcement. A few patterns that come up repeatedly:

Ask for the process. Outlines, annotated drafts and a short reflection on what changed between versions are hard to fabricate convincingly and useful to read regardless.

Anchor to specifics only they have. A seminar discussion, a local case, a dataset from your own module. Generic prompts get generic answers, from people and models alike.

Write the policy precisely. State which uses are fine, which need disclosure, and which are not permitted — and say explicitly whether spellcheck and grammar tools count. Most disputes trace back to a policy that said "no AI" and left everyone guessing.

If you want a second opinion on a score

Our detector is free for 10,000 words a month, up to 1,500 in a single scan, and reports sentence by sentence rather than as one document number — so you can see where a tool's suspicion concentrates instead of taking a percentage on trust. It returns human, AI or mixed, with the emphasis on mixed, because most real documents are.

It is subject to every limitation on this page. We sell a humanizer as well, which is a conflict of interest worth stating plainly: our editorial policy sets out how we handle it. HumanFlow cannot tell you whether a student used AI, and no tool that claims otherwise should be believed.

Run a free scan →

Common questions

How accurate are AI detectors for teachers?
Accurate enough to be a signal, not accurate enough to be evidence. The Liang study found seven detectors averaged a 61.22% false positive rate on TOEFL essays written by real people. Turnitin claims under 1%, but only for documents where at least 20% of the text is flagged — below that it shows an asterisk instead of a score.
Can I fail a student based on an AI detection score?
No detector output supports that on its own, and most vendors say so themselves. A score is a prompt to look further — at drafts, version history and the student's prior work. Institutions set their own policies, and every one we are aware of treats the score as an input to a human judgement rather than a finding.
Which students get falsely flagged most often?
Non-native English speakers, by a wide margin in the published research. Also students taught rigid essay structures, technical and scientific writers, and heavy self-editors. What those have in common is prose that is unusually even and conventional — which is exactly what detectors read as machine-like.
What should I do before accusing a student?
Work through the checklist on this page. In short: check the score is above the tool's own reliability threshold, look for process evidence, consider whether the student is in a known false-positive group, and open the conversation as a question rather than a verdict.
Do different AI detectors agree with each other?
Often not. The same passage can score very differently across tools, because each was trained differently and each sets its own thresholds. If two detectors disagree about a student's essay, that disagreement is information about the detectors rather than about the student.
Is there a free AI detector I can use as a teacher?
Several, ours included — 10,000 words a month free, up to 1,500 in a single scan, with a sentence-level readout. Grammarly's is free with no stated cap. Use any of them as a starting point for a conversation, never as its conclusion.
What is a reasonable AI policy to set for an assignment?
Specific beats strict. Say which uses are allowed, which require disclosure, and which are prohibited, and give a sentence of example wording for disclosure. Most disputes we hear about come from policies that said 'no AI' without defining whether that covered spellcheck, grammar suggestions or research.

Sources

  1. 1.GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press), 2023
  2. 2.AI detection tools falsely accuse international students of cheating The Markup, 2023
  3. 3.Understanding false positives within our AI writing detection capabilities Turnitin, 2023
  4. 4.New research: Turnitin's AI detector shows no statistically significant bias against English language learners (Turnitin-funded) Turnitin, 2024
  5. 5.Turnitin's AI writing detection capabilities FAQs Turnitin Guides, 2026
  6. 6.OpenAI scuttles AI-written text detector over 'low rate of accuracy' TechCrunch, 2023
  7. 7.Contra generative AI detection in higher education assessments arXiv, 2023

Related

Are AI detectors accurate? · Why human writing gets flagged · How detectors work · Turnitin AI detection in full · The student-facing version of this page