Why are non-native English writers flagged more often?
Because detectors measure how predictable text is, and writing carefully in a second language produces predictable text — in peer-reviewed testing, seven detectors flagged 61.22% of TOEFL essays by non-native speakers as AI-generated while classifying US 8th-grade essays almost perfectly.
Last reviewed 15 August 2026 · The HumanFlow team
The mechanism is not bias in the everyday sense of a rule written against anyone. It follows from what the tools measure. A detector scores text by how closely each word follows from the last, and a writer working in a second language sensibly stays with constructions and vocabulary they are confident in. That is competence under pressure, and it produces exactly the low-surprise prose the detector was built to flag.
The Liang et al. study in Patterns is the measurement people cite, and the size of the gap is what makes it damning: 89 of 91 TOEFL essays tripped at least one of the seven tools, while essays by native-speaking US eighth-graders were classified correctly at a much higher rate. Same tools, same settings, radically different error rates by population.
The study also tested the reverse. Prompting a language model to rewrite the TOEFL essays with richer vocabulary reduced the misclassification substantially — which confirms the mechanism, since making the text less predictable made the flag go away without changing who wrote it.
Turnitin has published research reporting no statistically significant bias against English language learners in its own detector. That is a vendor testing its own product, and it is a real data point rather than a dismissable one; it is also a different tool from the seven tested in the Patterns paper.
When this answer changes
Detectors are retrained regularly, and a finding measured on the tools available in 2023 does not automatically describe the same product today. What has not changed is the underlying mechanism, because predictability is still what these systems measure.
Institutions that have disabled AI detection have frequently cited disparate impact on international students among their reasons, so the practical exposure varies with where you study.
Where to go next
- The false positive problem in full — Who else it lands on, and what the published rates actually cover.
- Writing in English as a second language — Practical ground for writers in this position, without a bypass promise.
- If it has already happened — What to do first, and what evidence carries weight.
- Detection outside English — Thinner testing, uneven training data, and the same population exposed again.
Sources
One of our direct answers.