humanflow
Turnitin · The HumanFlow team · 13 min read

Turnitin AI Score Meaning: What the Number Actually Tells You

Turnitin's AI score is the share of prose its model thinks is AI-written — a probability estimate, not proof. Here's what 0%, *, 20%, and 100% mean.

Turnitin's AI score is the percentage of qualifying prose in a document that Turnitin's model classifies as likely AI-generated. It is a statistical estimate, not evidence — Turnitin tells instructors directly that the score alone should never be the basis for an accusation. A 40% doesn't mean you're 40% guilty, and a 0% doesn't mean you're clean.

That's the short version. The longer version matters, because this single number now shapes conversations in academic integrity offices around the world, and most of the people in those conversations — on both sides of the desk — misread what it says.

One scope note first. Turnitin's report contains two numbers, and this article covers only the newer one. The similarity score measures text overlap with existing sources and comes with checkable evidence. The AI score is a model's opinion about how your prose was produced and comes with no evidence at all — nothing to compare against, no source to inspect. If your issue is the overlap number, you want our companion guide to what similarity score is too high. Everything below is about the AI indicator, launched April 4, 2023, inside the existing Similarity Report.

What the number is, mechanically

Turnitin doesn't read your paper and form an impression. Its detector breaks qualifying text into overlapping segments — chunks of continuous prose — and runs each through a classifier trained to distinguish human-typical writing from machine-typical writing. Each segment gets a probability. Segments that cross Turnitin's internal threshold count as AI-written; the score you see is roughly the share of the document's prose those flagged segments represent.

"Qualifying" is doing real work in that sentence. The system needs roughly 300 words of continuous prose to run at all. It was built and validated primarily on English long-form writing. Bullet lists, code, equations, tables, and references don't classify well and are largely set aside. So the score is a percentage of the prose the model felt able to judge, not of the whole file — one reason two documents of very different lengths can carry the same score while meaning quite different things.

What is the classifier actually looking at? The same two families of signal as nearly every detector: perplexity — how predictable each next word is, since language models generate text by picking probable next words, making their output measurably smoother than human prose — and burstiness, the variation in sentence length and structure that humans produce naturally and models historically flatten out. The mechanics are covered in depth in how AI detectors work, but one sentence carries the essential point: the detector measures how machine-typical your text is, not who wrote it. Those usually correlate. When they don't, you get a false positive — or a false negative.

The score-interpretation table

Here is what each range means, and — just as important — what it doesn't.

Score shownWhat it meansWhat it does NOT mean
0%The model classified no qualifying segments as AI-likeProof the paper is human-written. Turnitin's product chief has said the system deliberately leaves roughly 15% of AI text unflagged to reduce false accusations (BestColleges, April 2023), and lightly edited AI prose slips through routinely
* (asterisk)Between 1% and 19% of prose was flagged — Turnitin suppresses the exact number because scores in this range "are less reliable" by its own admissionA hidden accusation. The asterisk exists precisely because Turnitin doesn't stand behind low-range numbers; treating * as "some AI detected" overreads it
20–39%A meaningful minority of segments read as machine-typical. This is the floor above which Turnitin's accuracy claims (98% accuracy, <1% false positives) even applyA finding that a fifth to a third of the paper was AI-written. Segment classification is lumpy; grammar tools, formulaic sections, and non-native phrasing patterns can contribute
40–79%The model found sustained machine-typical prose across much of the documentCertainty. Turnitin instructs institutions that even high scores warrant human review, not automatic penalty
80–100%Most or all qualifying prose classified as AI-like — the pattern raw chatbot output typically producesA confession. Human writers with flat, formulaic styles — especially non-native English writers — have hit high scores on genuinely original work

Two rows deserve expansion, because they're the two most misread.

The asterisk is not a euphemism. When Turnitin replaced 1–19% scores with a * in 2023, it was publicly conceding that its own model produces too many false positives in the low range to justify showing a number. That's genuinely responsible engineering — a vendor voluntarily hiding output it can't defend — and it deserves to be read as designed. A * means "the model saw something but we don't trust it enough to quantify." Any instructor treating an asterisk as grounds for suspicion is using the tool against its manufacturer's explicit instructions.

The 20% line is where the accuracy claims begin, not where guilt begins. Turnitin's headline figures — 98% accuracy, under 1% false positives — apply only to documents where more than 20% of the text is flagged. Below that line, Turnitin makes no comparable claim. Above it, the claims are the company's own, from its own testing, and independent checks haven't always matched them: The Washington Post's April 2023 test watched Turnitin flag an innocent student's writing and stumble on mixed human-AI drafts, which Turnitin acknowledged are its hard case. Blended documents — a human draft polished with AI, an AI draft rewritten by a human — are exactly where segment-level classification gets noisy, and they're also the most common real-world scenario.

Why the same paper can score differently over time

Students occasionally run an old paper through a checker twice and get different numbers, then conclude the whole system is random. It isn't random, but it isn't stable either, and the reasons are worth knowing.

The detector is a model, and Turnitin updates it. The version scoring papers today is not the version that launched in April 2023 — Turnitin added AI paraphrasing detection in July 2024 and AI bypasser detection in August 2025, both English-only, and retrains as the models it's trying to detect evolve. A paper scored under one model version can legitimately score differently under the next.

The target also moves. GPT-4-era prose has different statistical fingerprints than the output of newer models; training data that made the classifier sharp against 2023 chatbots ages as generation styles change. And segmentation itself introduces sensitivity: because the score aggregates threshold decisions over chunks of text, small changes — a reformatted reference list, a converted file format, a few edited sentences shifting segment boundaries — can move borderline segments across the line and swing the headline number more than the underlying edit would suggest.

The practical upshot cuts both ways. A score is a snapshot of one model version's opinion on one parse of one file. That's a reason for institutions to hesitate before building high-stakes decisions on it, and a reason for students to stop treating any single scan — Turnitin's or a third-party checker's — as a stable fact about their document.

What Turnitin says the score is for

Turnitin's own instructor guidance is more modest than the way the score is often used. The company frames the AI indicator as information to support a conversation — a signal that a document deserves a closer look — and states that it should not be the sole basis for an academic integrity action. The percentage identifies how much text the model flagged, and the report highlights which passages, so an instructor can read those passages, compare them with the student's other work, and make a human judgment.

Take that framing seriously, because it's load-bearing. Turnitin makes a strong accuracy claim (98%, above the 20% line) while simultaneously insisting the output isn't proof. Both halves are consistent once you do the base-rate arithmetic: even a 1% false-positive rate, applied across an institution's entire submission volume, wrongly flags real students in absolute numbers that add up fast. That math is precisely why Vanderbilt University disabled the indicator in August 2023, publishing its reasoning rather than quietly flipping a switch. Turnitin screened over 200 million papers in the indicator's first year; about 11% showed 20%-plus AI writing. At that scale, "rare" errors are a daily occurrence somewhere.

And the error risk isn't evenly distributed. The best-documented finding in this field is Liang et al.'s 2023 study in Patterns: seven GPT detectors evaluated on 91 essays genuinely written by non-native English speakers falsely flagged an average of 61.22% of them, and 89 of the 91 were flagged by at least one detector — while the same detectors were near-perfect on essays by native-speaking US eighth graders. Turnitin was not among the seven tested, so don't quote that number at Turnitin; but the study targets the exact statistical approach (predictability and uniformity of prose) that Turnitin's classifier shares. Writers whose English is fluent but formulaic — non-native speakers, students drilled in rigid essay structures, technical writers, some neurodivergent writers — produce human text with machine-typical statistics. The full picture is in our guide to AI detector false positives.

Reading your own score honestly

If you're a student staring at a number, here's the honest read for each situation.

You didn't use AI and got flagged anyway. It happens, it's documented, and it's survivable. Your defense isn't rhetoric; it's process evidence. Version history in Google Docs or Word, notes, outlines, earlier drafts, and your track record of in-class writing are worth more than any counter-scan. Ask to see which passages were flagged and offer to discuss them — genuine authors can talk about their own sentences in a way that's hard to fake.

You used AI within your course's rules — brainstorming, grammar cleanup, or whatever your syllabus permits. Disclose it as required. A score plus a disclosure is a non-event; a score plus a discovered omission looks like concealment even when it wasn't.

You used AI against your course's rules. No interpretation of the score changes what that is. If AI assistance is banned and you used it, the issue is the violation, not the detection — and editing the output to lower a number is a second choice on top of the first, not a fix for it.

You wrote with AI as a drafting partner in a course that allows it, and you're worried the statistical residue of machine phrasing misrepresents work that is substantively yours. This is the legitimate middle ground, and it's where editing in your own voice — restructuring arguments, adding your examples, breaking the uniform rhythm — matters. Tools exist for this; HumanFlow's AI humanizer is one, and it doesn't promise to beat Turnitin or any detector, because nobody can honestly promise that. What rewriting-in-your-voice actually does is make the text more yours, which is the only version of this that's defensible anyway.

The score in context

Zoom out and the AI score is one artifact of a hard problem: distinguishing human from machine prose is statistically possible on average and unreliable in individual cases. OpenAI — the company with the best possible view of its own outputs — shipped a text classifier that caught only 26% of AI writing while falsely flagging 9% of human writing, then retired it in July 2023 citing low accuracy. Turnitin's detector is meaningfully better than that — it scored highest of the 14 tools in the largest peer-reviewed multi-tool test, though that test still placed every tool below 80% accuracy. Both things are true: raw AI output usually gets caught, and individual flags on individual humans are wrong often enough that no fair process can treat the number as a verdict.

Which is why the most useful question about your score usually isn't "what does the number mean" but "what will a human do with it." Instructors range from careful to careless in how they act on a flag, and knowing the difference tells you far more about your actual risk than the percentage does. That's the subject of the next post in this series: how professors actually use Turnitin's AI report. For the full map of Turnitin's detection system — similarity, AI, accuracy claims, institutional pushback — start at the Turnitin detection hub.

FAQ

What does the Turnitin AI score actually measure? It's the percentage of a document's qualifying prose that Turnitin's classifier judged likely AI-generated, based on statistical patterns like word predictability and sentence uniformity. It measures how machine-typical the text is, not who wrote it. Turnitin says it should support human review, never replace it.

What does the asterisk (*) on a Turnitin AI score mean? It means the model flagged between 1% and 19% of the prose, and Turnitin suppresses the exact figure because it considers scores in that range unreliable. It is explicitly not a finding of AI use. Treating an asterisk as suspicion contradicts Turnitin's own design intent.

Does a 100% Turnitin AI score prove a paper was AI-written? No score proves authorship; 100% means essentially all qualifying prose classified as machine-typical, which is the pattern raw chatbot output produces — but human writers with flat, formulaic styles have been wrongly flagged at high percentages. Turnitin itself directs instructors to treat any score as the start of a review, not its conclusion.

Can Turnitin's AI score be wrong? Yes, in both directions. Turnitin claims under 1% false positives, but only for documents with more than 20% flagged, and its chief product officer has said it lets "probably 15%" of AI writing go by to hold false positives under 1%. Independent research on similar detectors found heavy bias against non-native English writers (Liang et al., Patterns, 2023).

Why did my paper get a different AI score the second time? Turnitin updates its detection model, and the score also depends on how the document is segmented into classifiable chunks. Model version changes, file format differences, and small edits near segment boundaries can all move the number. A score is one model version's snapshot, not a stable property of your text.

Is a 20% AI score bad? Twenty percent is the threshold above which Turnitin's accuracy claims apply at all, so it's where instructors are told the number becomes meaningful — but it is still an estimate, not evidence. Context decides everything: course rules on AI, which passages were flagged, and how your explanation matches your drafting history.

Do students get to see their Turnitin AI score? By default, no — Turnitin shows the AI indicator to instructors and administrators, not students. Some instructors share it voluntarily during a review conversation. If you've been flagged, asking to see the flagged passages together is a reasonable and normal request.

Is the AI score the same as the similarity score? No, and confusing them is the most common Turnitin misunderstanding. The similarity score matches your text against existing sources and shows checkable evidence; the AI score is a probability estimate with no source behind it. Our similarity score guide covers the other number.

Key facts

  • Turnitin launched its AI writing indicator on April 4, 2023, inside the existing Similarity Report (Turnitin).
  • Turnitin's accuracy claims — 98% accuracy, <1% false positive rate — apply only to documents where more than 20% of text is flagged as AI (Turnitin AI writing FAQ).
  • Scores of 1–19% display as an asterisk, not a number: Turnitin's own acknowledgment that low-range scores are unreliable (Turnitin).
  • The detector requires roughly 300 words of continuous prose and was built primarily for long-form English writing (Turnitin).
  • Turnitin's chief product officer has said the system deliberately leaves about 15% of AI text unflagged to reduce false accusations (BestColleges interview) (BestColleges, April 2023).
  • In the indicator's first year, Turnitin screened 200M+ papers; ~11% showed ≥20% AI writing and ~3% were ≥80% AI (Turnitin first-anniversary release, April 2024).
  • OpenAI's own AI text classifier caught only 26% of AI writing while falsely flagging 9% of human writing, and was retired in July 2023 (OpenAI).
  • Vanderbilt University disabled Turnitin's AI indicator in August 2023, publishing false-positive math at institutional scale as its reasoning (Vanderbilt statement).

Sources

  1. Turnitin — AI Writing Detection FAQ and transparency documentation (launch date, 20% threshold, asterisk policy, ~300-word minimum, instructor guidance).
  2. Turnitin — first-anniversary AI detection data release, April 2024.
  3. Liang, W., et al. "GPT detectors are biased against non-native English writers." Patterns (Cell Press), 2023.
  4. OpenAI — announcement retiring the AI text classifier, July 2023.
  5. Vanderbilt University — statement on disabling Turnitin's AI detection, August 2023.
  6. Fowler, G. "We tested a new ChatGPT-detector for teachers. It flagged an innocent student." The Washington Post, April 2023.
  7. BestColleges — "We Tested Turnitin's New AI Detector," April 21, 2023 (Annie Chechitelli on the 85%/15% trade-off).
All postsPublished by The HumanFlow team