The most-cited AI detector bias study — Liang et al., published in Patterns in 2023 — ran 91 human-written TOEFL essays through seven detectors and found they were falsely flagged as AI 61.22% of the time on average, while essays by native-speaking US eighth-graders passed almost untouched. One paper, and it rewired how institutions treat detection scores.
Three years on, "the Stanford study" gets invoked constantly and read rarely. It's cited by universities disabling detectors, by students appealing accusations, by vendors claiming they've fixed the problem, and — full disclosure — by companies like ours. We build a detector and humanizer ourselves; here's our editorial policy, and this post is exactly the kind of source-checking it commits us to. What follows is the complete anatomy of the paper: the design, every headline number, the mechanism, the prompt-engineering experiment buried in the same pages that almost nobody mentions, the honest criticisms, and what has and hasn't changed since.
The paper and the people
The full citation: Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, and James Zou, "GPT detectors are biased against non-native English writers," Patterns (Cell Press), volume 4, July 2023. The team came from Stanford — Zou's lab works on trustworthy machine learning — and the paper went out as open access, which is part of why it traveled so far, so fast. A preprint hit arXiv in April 2023, days after Turnitin switched on its AI indicator for the world's universities. The timing was not planned. It was perfect anyway.
The research question was narrow and sharp: do GPT detectors treat native and non-native English writers the same? Not "are detectors accurate" — a dozen papers were already circling that. This one asked who pays when they're wrong.
The design: seven detectors, two piles of essays
The setup is almost elegant in its simplicity, which is another reason it stuck.
The authors collected 91 human-written TOEFL essays from a Chinese educational forum — practice essays for the Test of English as a Foreign Language, written by real students, posted before ChatGPT existed in any form that could have contaminated them. As a control, they took 88 essays by US eighth-graders from the ASAP dataset, a public corpus of American student writing originally assembled for automated-scoring research funded by the Hewlett Foundation. Both piles were verifiably human. The only meaningful difference between them: one was written by non-native English speakers, one by native-speaking children.
Then they fed both piles to seven widely used detectors, as of early 2023: Originality.ai, Quil.org, Sapling, OpenAI's own AI text classifier, Crossplag, GPTZero, and ZeroGPT. Note who's absent: Turnitin. The most consequential academic detector wasn't tested — its access model made that impractical — and any honest summary has to say so. What the paper shows is a pattern across the detection approach Turnitin's category shares, not a measurement of Turnitin itself. (Our Turnitin hub covers what Turnitin does publish about its own error rates.)
Every essay in both piles was human-written. So every single flag, from any detector, was a false positive by construction. That's the beauty of the design: no arguments about ground truth, no "well, maybe the student really did use AI." The students in question wrote these essays before the tools existed.
The results, in full
Here is what came back, and it remains startling to type out.
| Measurement | TOEFL essays (non-native, n=91) | US 8th-grade essays (native, n=88) |
|---|---|---|
| Average false positive rate across 7 detectors | 61.22% | ~5.19% (most detectors near-perfect) |
| Flagged as AI by at least one detector | 89 of 91 (97.8%) | — |
| Flagged as AI by all seven detectors unanimously | 18 of 91 (19.8%) | — |
Sit with the middle row for a second. Ninety-seven point eight percent of these human writers would have been accused by somebody. If a student's essay were checked against the field of available detectors — which is exactly what a suspicious instructor shopping between free tools does — nearly every one of these 91 real people would have tripped an alarm somewhere. And roughly one in five essays was declared AI-generated by every single detector, unanimously, seven votes to zero, about text written by humans before the technology existed.
Meanwhile the eighth-graders walked through. The paper reports the detectors evaluated the native-speaker control "near-perfectly" — the average false positive rate on the ASAP essays worked out to about 5.19%, with a majority of the seven detectors flagging essentially none of them.
The gap between 61.22% and 5.19% is the study. Detectors weren't randomly wrong. They were wrong at one specific population, over and over, in the same direction.
The mechanism: why fluent-but-plain reads as machine
The authors didn't stop at the scandal; they explained it, and the explanation is the part worth carrying with you.
Most 2023-era detectors leaned on perplexity — roughly, how surprised a language model is by each next word. Machine text is built by picking probable words, so it's low-perplexity by construction. The detectors' wager was that human text is reliably lumpier. (Our how-detectors-work explainer unpacks perplexity and its cousin, burstiness, properly.)
Now think about what a non-native speaker's exam English looks like. TOEFL prep explicitly rewards safe vocabulary, standard transitions, and conventional sentence templates. A writer working in a second language draws from a smaller lexical pool and takes fewer syntactic risks — rationally, because risks cost points. The result is fluent, correct, predictable prose. Low perplexity. Machine-typical, to a machine.
The paper quantified it: the 18 essays unanimously misclassified by all seven detectors showed significantly lower text perplexity than the rest (P = 9.74 × 10⁻⁵), and the TOEFL essays as a group showed lower lexical richness, lexical diversity, and syntactic complexity than the eighth-graders'. The detectors, the authors concluded, were effectively penalizing "writers with limited linguistic expressions." The bias isn't a bug bolted onto the method. It is the method, applied to people whose honest writing happens to share statistical texture with model output. The same logic endangers other predictable-by-training writers — students drilled on rigid five-paragraph structures, technical writers, and, as later reporting documented, some neurodivergent writers. We map that fuller risk ordering on our false positives page.
The twist almost everyone skips: the prompt experiment
The second half of the paper is less quoted and more damning, because it closes the trap from both sides.
Side one: fixing the humans. The authors asked ChatGPT to revise the TOEFL essays with a simple instruction — "Enhance the word choices to sound more like that of a native speaker." After this machine-assisted polish, the average false positive rate on the (still fundamentally human) essays collapsed from 61.22% to 11.77%. Read that again: running human writing through an AI made detectors more confident it was human. The self-defense move available to the wrongly accused was to use the very tool they were accused of using.
Side two: hiding the machines. They then generated US college admission essays with ChatGPT-3.5 and confirmed detectors initially caught them well — up to 100% detection on the raw output. One prompt changed everything: "Elevate the provided text by employing literary language." Detection fell from 100% to 13%. A parallel experiment on GPT-generated scientific abstracts, using "Elevate the provided text by employing advanced technical language," dropped detection from a maximum of 68% to a maximum of 28%.
So within one paper: the detectors flagged innocent humans more than half the time, and a one-line prompt let actual machine text sail past them. High false positives and trivially achievable false negatives, from the same perplexity foundation. Whatever these tools were measuring, it wasn't authorship. It was fanciness of phrasing.
Being plain about our own position here: this finding is precisely why we refuse to sell "undetectable" as a promise. The paper proves detectors can be fooled — and also that they misfire on innocents, which means an arms race of evasion and detection mostly manufactures collateral damage. Our AI humanizer exists for editing machine-assisted drafts into your own voice where AI use is permitted; it doesn't promise to beat any detector, because nobody can honestly promise that, and this study is Exhibit A for why.
The honest criticisms
A paper this widely weaponized deserves adversarial reading. Here is the strongest case against over-reading it — offered seriously, not as a wink.
The sample is small and narrow. Ninety-one essays, one source forum, predominantly one first-language background (Chinese). That's a demonstration, not a census. The 61.22% figure is an average over seven specific tools on one specific corpus; it was never a universal constant, and quoting it as "AI detectors are wrong 61% of the time" — which happens weekly on social media — misstates the paper.
The detectors were early-2023 vintage. This is the big one. The seven tools tested included OpenAI's classifier, which OpenAI itself retired in July 2023 for low accuracy — it caught only 26% of AI text and false-flagged 9% of human text by OpenAI's own accounting. Several others have been retrained repeatedly since. Vendors are right when they say the paper describes their 2023 product, not necessarily their 2026 one. Pangram, notably, markets its abandonment of perplexity-style detection as the fix for exactly this failure mode, and newer Copyleaks and Turnitin materials advertise sub-1% false positive claims. Whether those claims survive independent replication is a different question — see below — but "the field moved" is a fair objection.
Corpus subtleties. TOEFL practice essays scraped from a forum may skew toward polished, template-following exemplars — the kind students post proudly — which could inflate the machine-typical signal. And the control group differs from the treatment group in more than nativeness: age, genre, and stakes all differ between a 13-year-old's classroom essay and an adult's exam prep. A cleaner design would match adult native-speaker exam essays against adult non-native ones. The original contrast is not as surgically controlled as the headline suggests, and we are not going to paper over that by gesturing at replications we cannot name.
Thresholds were vendor defaults. Each detector's yes/no flag depends on a tunable cutoff. A vendor could argue their tool, at a stricter institutional threshold, would have flagged fewer essays. True — though this cuts both ways, since real accusers used the same defaults the researchers did.
None of these criticisms rescues the detectors. A method that fails this badly, this directionally, on any plausible corpus of real human writing has a validity problem, whatever its version number. But precision about what the paper does and doesn't show is the difference between citing evidence and waving a talisman.
What changed since — and what didn't
Changed: The worst detector in the study is dead; OpenAI killed its classifier within months. Major vendors moved beyond naive perplexity toward trained classifiers on large paired corpora, and several now publish per-language accuracy claims and non-native-English evaluations of their own. Turnitin — never in the study — built visible guardrails: its 98%-accuracy and <1%-false-positive claims apply only to documents flagged over 20% AI, scores of 1–19% display as an asterisk rather than a number, and it requires roughly 300 words of prose to score at all. Those are real concessions to the false-positive problem, and they deserve to be called responsible engineering. (We track the current state of vendor claims and independent evidence on our accuracy hub.)
Didn't change: No detector vendor has published an independent, adversarial, non-native-speaker evaluation at the standard this paper set, run by someone they don't pay. The RAID benchmark (ACL 2024), the field's main independent stress test, found detectors still "easily fooled by adversarial attacks... and unseen generative models" despite 99%+ marketing claims — the same brittleness Liang et al. exposed with one sentence of prompting. And the underlying statistical reality is untouched: any detector that scores predictability will always sit closest to the error boundary for writers whose honest prose is predictable. You can shrink that bias with better training data. You cannot define it away.
Why one paper reshaped policy
Institutions don't usually move on a single study. This one had three properties that made it unignorable.
It arrived at the exact moment of decision — spring 2023, as Turnitin's indicator switched on by default across thousands of institutions and administrators were deciding what to do with the new number. It named a legally sensitive victim class: universities can absorb "detectors are sometimes wrong," but "detectors systematically accuse international students" reads like a discrimination complaint waiting for a plaintiff. And it was simple enough to survive a committee meeting: seven detectors, 91 human essays, 61% falsely flagged. No ROC curves required.
Vanderbilt University disabled Turnitin's AI indicator in August 2023, publishing arithmetic rather than opinion: it had sent 75,000 papers to Turnitin in 2022, and at the vendor's own stated error rate "around 750 student papers could have been incorrectly labeled". Other large institutions followed by disabling the score or demoting it to advisory-only. Guidance documents across higher education began instructing faculty that a detector score alone is not evidence — phrasing that traces, directly or through intermediaries, to this paper's findings. The study also became standard ammunition in student appeals, for good reason: it is peer-reviewed, quantified proof that innocent writing gets flagged, concentrated on the population least equipped to fight back. Our companion posts on enterprise deployments and on running your own detector test both inherit its central lesson — never accept an accuracy claim without asking on whose writing.
The wry footnote to all of this: the paper's most practical finding for a wrongly accused student — that one ChatGPT polish pass makes human writing look more human to the machines — is the one finding no university guidance document has ever figured out what to do with.
FAQ
What is the Stanford TOEFL AI detector study? It's "GPT detectors are biased against non-native English writers" by Liang, Yuksekgonul, Mao, Wu, and Zou, published in Patterns (Cell Press) in July 2023. It tested seven AI detectors on 91 human-written TOEFL essays and 88 essays by US eighth-graders, finding a 61.22% average false positive rate on the non-native writers versus near-perfect performance on the native-speaker control.
Which detectors did the study test? Originality.ai, Quil.org, Sapling, OpenAI's AI text classifier, Crossplag, GPTZero, and ZeroGPT — all as they existed in early 2023. Turnitin was not among them, though it uses the same broad statistical family of detection.
Why do AI detectors flag non-native English writers? Because most detectors score predictability (perplexity). Writers working in a second language rationally use safer vocabulary and more conventional structures, producing lower-perplexity text that statistically resembles machine output. The study confirmed unanimously misclassified essays had significantly lower perplexity (P = 9.74E-05).
Is the study still valid in 2026? Its specific numbers describe 2023-era detectors, several of which are retired or retrained — that criticism is fair. Its mechanism remains valid: predictability-based scoring structurally over-flags predictable human writers, and independent benchmarks like RAID (ACL 2024) still find detectors far less reliable than advertised. Treat 61.22% as historical evidence, not a current spec sheet.
Did the study show detectors can be beaten? Yes — that's the overlooked half. Prompting ChatGPT to "elevate the provided text by employing literary language" dropped detection of AI-generated college essays from 100% to 13%. The same paper documents both over-flagging of humans and easy evasion by machines.
Can I cite this study in an academic misconduct appeal? It's commonly cited in appeals as peer-reviewed evidence that detectors produce false positives, particularly against non-native speakers. It shows detectors can be wrong, not that one was wrong about you — pair it with your own evidence: drafts, version history, and writing samples. This isn't legal advice.
What was the eighth-grader control group for? To isolate nativeness. Both essay piles were verifiably human, so any flag was a false positive; the detectors' ~5% error on native-speaker essays versus 61.22% on non-native essays showed the failures were concentrated on one population rather than random.
Key facts
- Liang et al., Patterns (Cell Press), July 2023: seven detectors tested on 91 human-written TOEFL essays and 88 US eighth-grade essays (ASAP dataset).
- Average false positive rate on the TOEFL essays: 61.22%; on the native-speaker control: ~5.19%, with most detectors near-perfect.
- 89 of 91 TOEFL essays (97.8%) were flagged by at least one detector; 18 of 91 (19.8%) were flagged unanimously by all seven.
- Unanimously misclassified essays had significantly lower perplexity (P = 9.74E-05) — the mechanism is predictability scoring, not authorship evidence.
- The prompt "Enhance the word choices to sound more like that of a native speaker" cut the TOEFL false positive rate from 61.22% to 11.77%; "Elevate the provided text by employing literary language" cut detection of AI-generated essays from 100% to 13%.
- Detectors tested: Originality.ai, Quil.org, Sapling, OpenAI's classifier, Crossplag, GPTZero, ZeroGPT. Turnitin was not tested.
- OpenAI retired its own classifier in July 2023 (26% catch rate, 9% false positives); Vanderbilt disabled Turnitin's AI indicator in August 2023.
Sources
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. "GPT detectors are biased against non-native English writers," Patterns 4(7), Cell Press, 2023 (open access; arXiv:2304.02819). Numbers verified against the full text, August 2026.
- OpenAI — "New AI classifier for indicating AI-written text," January 2023, updated July 2023 with retirement notice.
- Vanderbilt University — "Guidance on AI detection and why we're disabling Turnitin's AI detector," August 2023.
- Turnitin — AI writing detection FAQ / transparency page (98% claim conditions, asterisk policy, ~300-word minimum).
- Dugan, L., et al. "RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors," ACL 2024 (arXiv:2405.07940).
- ASAP (Automated Student Assessment Prize) dataset — Hewlett Foundation, source of the eighth-grade control essays.
- Fowler, G. The Washington Post, April 2023 — Turnitin false-positive test (contextual).