Turnitin and Non-Native English Speakers: The False-Positive Problem Nobody Warned You About
Research found AI detectors falsely flagged 61% of non-native English essays. Why ESL writing trips Turnitin-style tools, and how to protect yourself.
If English isn't your first language, AI detectors are measurably more likely to flag your honest writing as machine-generated. The landmark study found detectors falsely flagged an average of 61.22% of human-written TOEFL essays — while judging native speakers' essays almost perfectly. This isn't your writing failing. It's the technology failing you.
That deserves saying clearly at the top, because most students who get flagged assume they did something wrong. You didn't. The bias is documented in a peer-reviewed journal, acknowledged by universities, and explainable down to the mathematics. This post covers the study everyone in this debate cites, the mechanism behind the bias, what you can actually do about it — habits that protect you without dumbing down a single sentence — and what your instructors owe you in return. It sits alongside our wider work on how AI detection actually works, and it pairs with our practical resource for ESL writers using AI tools.
The study that proved the bias: Liang et al., 2023
In 2023, a Stanford-led team — Weixin Liang, James Zou, and colleagues — published a paper in Patterns, a Cell Press journal, with a title that gave away the verdict: "GPT detectors are biased against non-native English writers." It has become the most-cited piece of evidence in every serious discussion of detector fairness, and for good reason. The design was simple and the results were not subtle.
The researchers took 91 essays written by real people for the TOEFL — the Test of English as a Foreign Language, taken by students applying to English-speaking universities. Every essay was human-written; that's the whole point of the TOEFL. They ran the essays through seven commercial GPT detectors. Then they did the same with essays written by native-speaking US 8th graders.
The results, in full:
| Finding | Number | What it means |
|---|---|---|
| Average false-flag rate across 7 detectors on TOEFL essays | 61.22% | A majority of these human essays were called AI, on average |
| Essays flagged by at least one detector | 89 of 91 | Nearly every non-native writer would be flagged somewhere |
| Essays flagged by ALL seven detectors | 18 of 91 | One in five was unanimously — and wrongly — condemned |
| Same detectors on native US 8th-grade essays | Near-perfect | The tools work fine — on native speakers |
Sit with the second row for a moment. Eighty-nine of ninety-one. If those 91 students each submitted their essay at a school using some AI detector, all but two could have faced a flag for writing they demonstrably produced themselves in a proctored testing environment. And the 18 essays flagged by all seven tools represent the nightmare scenario: a student who appeals one detector's verdict, gets re-checked with another "for fairness," and loses again — not because they cheated, but because every tool shares the same blind spot.
One precision matters here, and we'll state it the way the researchers would want. Turnitin was not among the seven detectors tested. Nobody can honestly claim Liang et al. measured Turnitin's false positive rate, and anyone who tells you "Turnitin flags 61% of ESL essays" is misquoting the study. What the study established is that the detection approach itself — the statistical method Turnitin's detector also rests on — systematically penalizes non-native writing. Turnitin says it built and validated its tool with false positives in mind, and its safeguards (a 20% reporting threshold, an asterisk over low scores) are real. Whether those safeguards fully close a gap this large has never been independently verified. Vanderbilt University cited this exact research when it disabled Turnitin's AI detector in August 2023.
Why your careful English reads as "machine-typical"
The bias isn't a bug someone forgot to fix. It falls straight out of how these detectors work, and understanding the mechanism is genuinely useful — it tells you what the tools are actually reacting to in your writing. (The full technical story is in our how detectors work guide.)
Detectors like Turnitin's measure two main statistical properties. Perplexity is how surprising each word is, given the words before it — essentially, how far the text strays from the most predictable phrasing. Burstiness is how much sentence length and structure vary across the document. Large language models write with low perplexity (they literally work by picking probable words) and low burstiness (an even, steady rhythm). So the detector's core assumption is: predictable and uniform equals machine.
Now look at what writing in a second language actually demands, and the collision becomes obvious.
When you write in a language you're still mastering, you make rational choices. You reach for vocabulary you're sure of — the common, well-worn word instead of the rare one you might misuse. That lowers perplexity. You build sentences on structures you know are grammatically safe, and you reuse those structures, because a reliable pattern beats an ambitious error when a grade is on the line. That lowers burstiness. You were probably taught academic templates explicitly — topic sentence, evidence, link back — and you apply them faithfully, because that's what earned marks in every English class you've taken. Uniformity again.
Every one of those choices is intelligent. Every one is what a diligent student should do in a second language. And every one pushes your statistical fingerprint toward the exact profile detectors were trained to call machine-generated. The Liang team demonstrated this directly: when they had a language model rewrite the TOEFL essays with more varied, idiomatic vocabulary, false flags dropped sharply. The detectors weren't detecting AI. They were detecting limited word choice — and punishing the writers who could least afford it.
There's a bitter irony underneath, and it's worth naming. The fluent, native-speaking student who writes with casual idiom and messy varied rhythm sails through. The international student who worked twice as hard to produce correct, disciplined English gets flagged. The detector penalizes effort expressed as caution. That's not a moral failing in your writing. It's a design limitation in the measurement.
The stakes are higher for you — and everyone should admit it
A false accusation is painful for any student. For international students, the same accusation carries extra weight that domestic classmates rarely see.
An academic-integrity finding can threaten a student visa, since enrollment status and good standing are often conditions of staying in the country. It can end a scholarship that took years to win. It can mean explaining a disciplinary record in a second language, through an unfamiliar process, without family nearby, in front of a panel that may know nothing about detector error rates. Some students face all of this over a starred score their instructor misread. None of that is hypothetical; university ombuds offices and international-student advisors have been handling exactly these cases since 2023.
We say this not to frighten you but because the stakes justify the preparation this post recommends. You are not being paranoid by keeping evidence of your writing process. You are being proportionate.
What you can do: protection without dumbing anything down
The goal here is emphatically not "write worse so the algorithm likes you." You've spent years building your English. Nothing below asks you to disguise it, simplify it, or make it less yours. There are two tracks: evidence habits, which are the strong protection, and writing habits, which help at the margins and make you a better writer anyway.
Evidence habits (do these starting today)
Write in software that keeps version history. Google Docs does this automatically; Microsoft Word does with AutoSave to OneDrive. A revision trail showing your essay growing over hours and days — sentences appearing, typos being fixed, paragraphs moving — is the closest thing to proof of authorship that exists. AI-generated text arrives in large pastes; human writing accretes. Let the software record the accretion.
Keep your intermediate mess. Outlines, brainstorm notes, annotated readings, earlier drafts, the paragraph you cut. Don't tidy your process folder until the course ends. Mess is evidence.
Draft in stages you can show. If you plan in your first language and then write in English — a completely legitimate method — keep those first-language notes too. They demonstrate a thinking process no chatbot produces.
Know your baseline. Run a few of your own past essays — things you wrote before ChatGPT existed, if you have them — through a detector such as our free AI detector, which shows sentence-level results. If your natural style scores high, you've learned something important before it costs you: you're in the elevated-risk group, your evidence habits matter more, and you have concrete examples for any future conversation with an instructor. One honesty note: no detector's output predicts another's, and we don't claim ours predicts Turnitin — the tools disagree with each other too often for anyone to promise that.
Writing habits (marginal help, real growth)
These are the same techniques writing teachers recommend for developing an authentic voice. They happen to also raise perplexity and burstiness — the honest way.
Vary your sentence lengths deliberately. After a long, complex sentence, drop in a short one. Like this. This single habit does more for burstiness than any vocabulary change, and it makes prose better in every language.
Anchor claims in specifics only you would choose. Your own examples, your country's context, a detail from this week's lecture, a source your instructor assigned. Generic support reads machine-typical; particular support reads human, because it is.
Let your perspective in. Where the assignment allows first person, use it. "In my experience in Lagos hospitals..." is a sentence no model would generate for you, and it's usually stronger writing too.
Don't over-sand your drafts. Fix real errors, but resist polishing every sentence into the same smooth shape. A slightly uneven, clearly-yours paragraph is both better prose and statistically more human than a uniformly buffed one.
Notice what's absent from this list: thesaurus games, deliberately broken grammar, "detector-proof" rewriting tricks. Those degrade your writing, betray your effort, and don't reliably work. And to be plain about our own product: HumanFlow's humanizer exists for editing AI-assisted drafts into your own voice where AI use is permitted — it doesn't promise to beat Turnitin or any detector, because nobody can honestly promise that, and if your course bans AI entirely, no tool makes disguising it acceptable.
What your instructors owe you
This section is for the teachers, TAs, and integrity officers reading — and students, you're welcome to hand it to yours.
Instructors owe ESL students knowledge of the base rates. Anyone using an AI detector on work by non-native English speakers should know the Liang et al. numbers and Vanderbilt's reasoning, the same way anyone prescribing a medical test should know its false-positive profile for the patient in front of them. A tool that is near-perfect on native speakers and wrong most of the time on TOEFL-level writers is not one tool; it's two, and the instructor should know which one they're using.
They owe a conversation before an accusation. Turnitin's own guidance says the score should start a dialogue, not conclude one. For an ESL student, that dialogue is also diagnostic in the student's favor: a ten-minute discussion of the paper's argument, in which the student explains their choices and sources, is far better evidence of authorship than any percentage. Students who wrote their essays can talk about them. It really is that simple most of the time.
They owe process-based alternatives. Scaffolded drafts, in-class writing samples for baseline comparison, oral defenses, annotated bibliographies — assessment designs that make authorship visible make detector scores nearly irrelevant. And they owe restraint with starred scores: a 1–19% asterisk carries, by Turnitin's own admission, no reliable meaning at all.
None of this asks instructors to ignore misconduct. Real AI-substituted work exists in volume — Turnitin flagged millions of heavily AI-written papers in the detector's first year — and pretending otherwise helps no one, least of all the honest students competing beside the cheaters. The ask is narrower: apply the tool's own documented limits, and apply extra care exactly where the tool is documented to fail. If a flag does land on you despite everything, our guide to why detectors flag human writing walks through the response step by step, and our summary of what the accuracy research actually shows gives you the numbers for the meeting.
The bigger picture
The detector-bias problem is not permanent physics; it's the current state of an immature technology. Turnitin restricts its accuracy claims to English and to documents over its reporting threshold, displays an asterisk over its least reliable range, and deliberately under-flags rather than over-accuse — imperfect measures, but real ones, and better than most competitors offer. Research pressure from papers like Liang et al. is exactly what pushed the industry toward those admissions. More pressure will help more.
Until then, your English is not the problem. Write in your voice, at your level, with your effort visible — and keep the receipts.
FAQ
Does Turnitin flag non-native English speakers more often? No published study has measured Turnitin's false-positive rate on ESL writing specifically. What's established, by Liang et al. in Patterns (2023), is that seven other detectors using the same statistical approach falsely flagged an average of 61.22% of human-written TOEFL essays while judging native speakers' essays near-perfectly. Turnitin says its safeguards address false positives; independent verification of that claim for ESL writing doesn't yet exist.
Why does my honest writing get flagged as AI? Detectors flag text that is statistically predictable (low perplexity) and structurally uniform (low burstiness) — and careful second-language writing tends to be both, because you sensibly choose safe vocabulary and reliable sentence patterns. The tools measure how machine-typical text is, not who wrote it. Disciplined, correct English written under constraint happens to resemble machine output statistically.
Should I use more complex words so detectors don't flag me? Not artificially — swapping in thesaurus words you wouldn't naturally use tends to weaken your writing and can introduce errors that cost real marks. The Liang study did show richer vocabulary reduces false flags, but the sustainable version of that is growth over time plus specific, personal content: your own examples, varied sentence lengths, your genuine perspective. Evidence habits (version history, drafts) protect you more than any vocabulary trick.
What's the single best protection against a false accusation? Version history. Write in Google Docs or Word with AutoSave, so every essay carries a timestamped record of being built gradually — the one form of evidence that's nearly impossible for a wrongful accusation to argue around. Keep outlines, notes, and earlier drafts as backup. Start this before you need it.
I've been flagged. What do I do right now? Don't panic and don't confess to something you didn't do. Ask to see the exact report and score, gather your version history and drafts, and offer to discuss the paper's content in person. Then read your institution's academic-integrity procedure so you know your response and appeal rights, and bring the Liang et al. study and Vanderbilt's August 2023 decision to the meeting as context.
Can I mention the research without sounding confrontational? Yes — frame it as context, not accusation. Something like: "I wrote this myself and can show my drafts. I'd also like to share a peer-reviewed study showing detectors falsely flag most TOEFL-level writing, because I think it's relevant to how we read this score." Instructors are far more receptive to evidence offered calmly than most students expect.
Is it safe for me to use Grammarly or similar tools? Basic grammar and spelling correction is generally accepted and unlikely to transform your statistical fingerprint on its own, but full-sentence rewriting features push your text toward machine phrasing — and some courses restrict even editing tools, so check your syllabus and ask when unsure. Whatever you use, keep your pre-correction draft. It's one more receipt.
Key facts
- Liang et al., Patterns (Cell Press), 2023: seven GPT detectors falsely flagged an average of 61.22% of 91 human-written TOEFL essays.
- 89 of 91 TOEFL essays were flagged by at least one detector; 18 of 91 were flagged by all seven (Liang et al., 2023).
- The same detectors were near-perfect on essays by native-speaking US 8th graders — the bias is specific to non-native writing (Liang et al., 2023).
- Turnitin was not among the seven detectors tested, but its detector uses the same perplexity/burstiness-based statistical approach.
- Vanderbilt University disabled Turnitin's AI detector in August 2023, citing false-positive math (~750 potential wrongful flags per 75,000 papers annually) and the research on bias against non-native speakers.
- Turnitin's detector was built and validated primarily on English, needs roughly 300 words of prose, and displays scores of 1–19% as an asterisk it deems unreliable (Turnitin).
- Non-native English speakers top the documented false-positive risk order, ahead of formula-trained students, technical writers, heavy self-editors, and neurodivergent writers.
Sources
- Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., & Zou, J. — "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023.
- Vanderbilt University, Brightspace blog — "Guidance on AI Detection and Why We're Disabling Turnitin's AI Detector," August 16, 2023.
- Turnitin — AI Writing Detection FAQ / transparency page (English-primary validation, 300-word minimum, asterisk policy, accuracy conditions).
- Turnitin — first-anniversary data release (200M+ papers screened, AI-writing prevalence), April 2024.
- Fowler, G. — Washington Post test of Turnitin's AI detector, April 2023.
- Inside Higher Ed — "Professors proceed with caution using AI-detection tools," February 9, 2024.