humanflow

Turnitin AI detection: what it measures, and what it proves

Turnitin's AI writing indicator launched on 4 April 2023. It reads your prose and estimates what share of it was machine-generated — nothing is matched against a database. Turnitin says it keeps false positives under 1% above a 20% threshold. Below that it shows an asterisk instead of a number, because its own testing found more errors there.

Last reviewed 16 August 2026 · The HumanFlow team

Two scores, two entirely different systems

Most confusion about Turnitin starts here. A Similarity Report and an AI writing score are produced by separate machinery and mean separate things, and students routinely read one as though it were the other.

The similarity score is the old one. It matches your text against Turnitin's repository of published work, web pages and previously submitted papers, and reports how much overlaps. A high number means your words appear somewhere else. It is a matching problem, and it is largely a solved one.

The AI writing indicator compares your text against nothing. It reads the prose and estimates how machine-typical the writing is, using signals like perplexity — how predictable each next word is — and burstiness, the variation in sentence length and structure. Human writing tends to be uneven. Model writing tends toward smooth and medium-everything.

The practical consequence: you can score 0% similarity and still be flagged for AI, or score 40% similarity on a properly quoted literature review and have no AI flag at all. The two numbers do not constrain each other. How detectors work covers the mechanism in full.

The figures Turnitin publishes

These are Turnitin's claims about Turnitin, sourced below. They are not independent findings, and this page treats them as what they are.

WhatTurnitin's stated figure
Launched4 April 2023
False positive targetUnder 1%, only where 20%+ of qualifying text is predicted AI
Scores 0–20%No score, no highlights — an asterisk, due to more false positives
Minimum length300 words of long-form prose (raised from 150)
Maximum length30,000 words
LanguagesEnglish, Spanish, Japanese
File types.docx, .pdf, .txt, .rtf — under 100MB
Papers reviewed65 million by July 2023; 200 million+ by April 2024
Flagged at 80%+ AI2.1 million of the first 65 million — 3.3%

The 20% line is the one worth understanding. Turnitin does not claim under 1% false positives on every document — it claims that rate for documents where at least a fifth of the qualifying text reads as machine-written. Below that, it declines to give a number at all. A great deal of writing about Turnitin quotes the sub-1% figure with the condition stripped off.

The false positive argument, both sides

This is the part of the subject where the evidence genuinely conflicts, and where most coverage picks a side and stops reading.

Against detector reliability: peer-reviewed work by Liang and colleagues found GPT detectors are biased against non-native English writers, reporting severe misclassification of TOEFL essays written by humans. The Markup documented international students falsely accused. OpenAI withdrew its own classifier in July 2023, citing a low rate of accuracy — the company with the most access to how these models write concluded it could not reliably detect them. Our accuracy page goes through that research in detail.

For Turnitin's position: Turnitin published research testing roughly 2,000 samples from English language learners, reporting a false positive rate of 0.014 for ELL writers against 0.013 for native English writers on documents meeting the 300-word minimum — a difference it describes as not statistically significant. That is a direct answer to the bias charge, and it is a real result.

The disclosure that matters: Turnitin compensated the researchers who conducted that study. That does not make the finding wrong, and vendor-funded research is normal across the industry. It does mean the strongest evidence for Turnitin's accuracy was paid for by Turnitin, while the strongest evidence against it was not. Hold both facts at once.

Note also what the two studies measure. Turnitin's figure applies to documents over 300 words in its own product. The Liang findings cover detectors generally, on shorter samples. They are not measuring the same thing, which is part of why both can be honestly reported.

The 2025 humanizer detection update

On 27 August 2025 Turnitin announced detection aimed specifically at humanizer tools — software built to strip machine fingerprints out of generated text. Chief product officer Annie Chechitelli said the company had "researched and identified the signals and patterns of leading humanizers and have trained our model to identify them," adding that "humanizers also leave a statistical trace that can be learned and detected."

The claimed false positive rate for the detection system remained under one percent. Jonathan Bailey, writing at Plagiarism Today, called the capability impressive while noting the claim "needs to be independently verified by third-party testing." As far as we know, that verification has not been published.

We sell a humanizer, so read the next sentence with that in mind. HumanFlow does not promise to beat Turnitin or any other detector — nobody honestly can — and we do not publish bypass rates, because a bypass rate is a snapshot of a moving target measured by the party selling the tool. How we handle writing about a market we compete in.

What a score does and does not establish

A percentage is an estimate produced by a statistical model. It is evidence in the same way a smoke alarm is evidence of fire: worth investigating, not sufficient on its own.

Turnitin's indicator is not designed to be the sole basis of a misconduct finding, and institutions set their own policies about what happens next. In practice, what carries weight in an appeal is process evidence — version history in Google Docs or Word, notes, outlines, earlier drafts, and a track record of graded work in your voice. Those beat a percentage in both directions.

If you have been flagged and did not use AI, the worst move is confessing to end the discomfort. It converts a defensible position into an admitted violation. Our false positives guide walks through the meeting itself.

Everything we have written on Turnitin

30 articles, grouped by what you are trying to find out.

Does Turnitin detect it?

The tool-by-tool answers, and what changes between them.

Reading your report

What the numbers mean before you react to them.

If you have been flagged

The research on false positives, and what actually helps.

Policy, ethics and citation

Where the lines sit, and who draws them.

Turnitin against other detectors

How it compares with the tools it is measured against.

More on Turnitin

Compared against

  • Turnitin vs Copyleaks Turnitin against Copyleaks on what independent testing found, who each is sold to, and what happens to the text you submit.
  • Compilatio vs Turnitin Turnitin against Compilatio on what independent testing found, who each is sold to, and what happens to the text you submit.
  • Turnitin vs Originality.ai Turnitin against Originality.ai on what independent testing found, who each is sold to, and what happens to the text you submit.
  • Turnitin vs Grammarly Authorship Turnitin against Grammarly Authorship on what independent testing found, who each is sold to, and what happens to the text you submit.
  • Scribbr vs Turnitin Turnitin against Scribbr on what independent testing found, who each is sold to, and what happens to the text you submit.
  • Pangram vs Turnitin Turnitin against Pangram on what independent testing found, who each is sold to, and what happens to the text you submit.

Check what a detector sees

HumanFlow's detector is not Turnitin and cannot tell you what Turnitin will say. What it gives you is a read on how machine-typical your draft looks, with the signals behind the score and how many sentences fall in each band — and, on a paid plan, which lines those were, usually the flat, evenly-weighted ones worth rewriting on their own merits. Free plan covers 10,000 detection words a month, up to 1,500 in a single scan.

Run a free scan →See plans

Two longer pieces on the same tool: can Turnitin be wrong? and what the similarity score counts, which is the row people most often confuse with the AI one. If a percentage has been put to you as a threshold, there is no safe Turnitin AI percentage.

Common questions

Does Turnitin detect AI writing?
Yes. Turnitin added an AI writing indicator on 4 April 2023, separate from its long-standing similarity score. It estimates what share of a document's qualifying text was machine-generated. It does not compare your work against a database to do this — it reads the prose itself and judges how statistically machine-typical it is.
What is a good Turnitin AI score?
There is no official threshold, and no score is a pass or a fail on its own. Turnitin attributes no score and no highlights between 0% and 20%, showing an asterisk instead, because its own testing found more false positives in that range. Above 20%, Turnitin says it aims to keep false positives under 1%.
Why does my report show an asterisk instead of a percentage?
Because the detected share fell between 0% and 20%. Turnitin suppresses the number in that band deliberately — it found a higher incidence of false positives at low percentages, so it shows an asterisk rather than a figure that would read as more certain than the evidence supports.
How many words does Turnitin need to check for AI?
At least 300 words of prose in long-form writing, up to a maximum of 30,000. Turnitin raised the minimum from 150 to 300 because accuracy improves with more text. Files must be .docx, .pdf, .txt or .rtf, under 100MB, in English, Spanish or Japanese.
Can Turnitin be wrong about AI writing?
Yes, and it says so itself by hiding scores below 20%. Independent research has found AI detectors misclassify human writing at meaningful rates, with peer-reviewed work reporting severe bias against non-native English writers. Turnitin's own funded study reports near-identical rates for both groups above 300 words. Both findings are on this page.
Does Turnitin detect AI humanizer tools?
It says it does. On 27 August 2025 Turnitin announced detection aimed at humanizer or 'bypasser' tools, with its chief product officer saying the company had identified the statistical signals those tools leave. As Jonathan Bailey noted at the time, that claim still awaits independent third-party verification.
Can I check my Turnitin AI score before submitting?
Not against Turnitin itself unless your institution enables it — the detector is sold to institutions, not students. Any third-party checker gives you a different tool's opinion, not Turnitin's. That is useful as a rough signal and worthless as a guarantee, and anyone selling it as the latter is misleading you.
Does a high AI score mean I will be penalised?
No. A score is an input to a human decision, not the decision. Institutions set their own policies and Turnitin states its indicator is not intended as the sole basis for an academic misconduct finding. Version history, drafts and notes carry more weight in an appeal than any percentage does.

Sources

  1. 1.Turnitin's AI writing detection capabilities FAQs Turnitin Guides, 2026
  2. 2.File requirements for an AI writing report Turnitin Guides, 2026
  3. 3.Understanding false positives within our AI writing detection capabilities Turnitin, 2023
  4. 4.New research: Turnitin's AI detector shows no statistically significant bias against English language learners (Turnitin-funded) Turnitin, 2024
  5. 5.Turnitin marks one year anniversary of its AI writing detector Turnitin (press release), 2024
  6. 6.Turnitin's AI detection feature reviews more than 65 million papers Turnitin (press release), 2023
  7. 7.Turnitin Launches Anti-AI Humanizer Feature Jonathan Bailey — Plagiarism Today, 2025
  8. 8.GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press), 2023
  9. 9.AI detection tools falsely accuse international students of cheating The Markup, 2023
  10. 10.OpenAI scuttles AI-written text detector over 'low rate of accuracy' TechCrunch, 2023

Related

Two definitions this page leans on, each with a worked example: AI writing score — what the percentage and the asterisk actually claim — and plagiarism, which is a separate finding on separate evidence and is routinely conflated with this one on the same report. The other number on that report is the similarity score, and only one of the two can show you its evidence. The 20% band above is a detection threshold— a line a vendor chooses, which means the same document can be flagged or suppressed without a word of it changing. And the single figure at the top of the report is document-level detection: it cannot tell you whether two paragraphs or twenty sentences produced it, which is the first thing worth asking.

AI detection hub · Are AI detectors accurate? · Why human writing gets flagged · How detectors work · Essays · Compare humanizers

Reading the number

An asterisk is not a low score, and 20% is a display threshold rather than a limit — what a score actually means.

What it will and will not read

The AI writing report has documented file limits, and they explain most missing scores: PDFs are accepted within a 300-word floor, and PowerPoint is not accepted at all.

The other detectors, reviewed

If your institution runs something other than Turnitin, or you have been flagged by a tool you had not heard of: Copyleaks · GPTZero · Originality.ai · ZeroGPT · Scribbr · Compilatio · SafeAssign