humanflow

What is a good AI detection score?

There isn’t one. No vendor publishes a threshold that means innocent, and the two biggest both say in their own documentation that their number should not decide anything on its own. If a percentage is being read as a pass mark, that is a choice someone made locally — and it is a choice you are allowed to question.

Last reviewed 18 August 2026 · The HumanFlow team

The single most useful fact: the asterisk

If your Turnitin report shows *% instead of a number, that is not a glitch and it is not a low score. It is Turnitin refusing to show you a figure it does not trust. In its own words:

“To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed.”

So the entire band from 1% to 19% is one undifferentiated “not reliable enough to report”. A 3% and a 19% look identical, because the vendor has decided neither number means anything on its own.

Why 20% is not a limit

20% is where Turnitin starts displaying a number. It is not where misconduct begins, and reading it as a cap gets the logic backwards.

The reason the line sits there is accuracy, not severity. Turnitin states a false-positive rate “under 1% for documents with over 20% of AI writing”, validated against more than 700,000 pre-ChatGPT academic papers before each model release. Note the qualifier: that claim is scoped to documents above the threshold. Below it, the company is not claiming that accuracy — which is precisely why it withholds the number.

A department that has adopted “20% or under” as a rule has borrowed a display setting and turned it into a disciplinary one.

What the number actually counts

Not how machine-like your writing sounds. Turnitin defines the percentage as “the amount of qualifying text within the submission that Turnitin’s AI writing detection model determines was likely generated by AI or likely generated and modified by an AI paraphraser or bypasser”.

Two things follow. First, it is a proportion of text, so a long quotation-heavy piece and a short one behave differently. Second, “qualifying text” means prose sentences in long-form writing — lists, bullet points and non-sentence structures are excluded. The denominator is not your whole document, which is why the percentage rarely matches anyone’s intuition about their own work.

GPTZero’s number is a different kind of number

This trips people up constantly, because both are shown as percentages and they are not comparable.

GPTZero returns a document classification — HUMAN_ONLY, MIXED or AI_ONLY — with a probability it describes as “the chance that the detector is correct in its classification”. That is confidence about the tool’s own verdict, not a share of your document. An 85% from GPTZero and an 85% from Turnitin are two unrelated measurements that happen to share a symbol.

GPTZero also publishes its own accuracy at high confidence: “99.1% of human articles are classified as human, and 98.4% of AI articles are classified as AI”. As with any vendor figure, that is measured on data the vendor chose — see what independent testing found, which is consistently less flattering than what vendors report.

Both vendors tell you not to use it alone

This is the part worth quoting back to whoever sent you the score, because it is not our opinion.

Turnitin: the percentage “should not be used as the sole basis for action or a definitive grading measure by instructors”, and the tool “provides data for the educators to make an informed decision based on their academic and institutional policies”.

GPTZero: “These results should not be used to punish students”, and “the sentence-level classification should not be solely used to indicate that an essay contains AI.”

A process resting on the number alone is not following the vendors’ guidance. It is going further than the companies that sell the software are willing to go.

If a score has been used against you

Three specific things to establish, in order: which tool produced it, what that tool’s number actually measures, and what evidence exists beyond it. An asterisk is not a score. A 20% is not a limit. A GPTZero probability is not a proportion.

Then the general guidance applies — what to do when you have been accused and how to prove you wrote it, where draft history does far more work than arguing about a percentage ever will.

Where we stand

We sell a detector, so we have an interest in you believing detection scores mean something. Everything above is quoted from the vendors’ own documentation and linked, and the honest summary is that a score is a weak signal with no defensible cutoff. We have published no accuracy figure for our own detector and will not before our benchmark produces one.

Related

Common questions

What is a good AI detection score?
There is no threshold that means innocent and none that means guilty, and no vendor publishes one. Turnitin says outright that the percentage 'should not be used as the sole basis for action or a definitive grading measure by instructors'. GPTZero says 'these results should not be used to punish students'. If a number is being treated as a pass mark, that is the institution's decision, not the tool's instruction.
What does an asterisk instead of a percentage mean in Turnitin?
It means AI was detected below 20% and Turnitin is deliberately withholding the number. In its words: 'To avoid potential incidence of false positives, no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed.' The asterisk is a statement that the reading is not reliable enough to show you.
Is 20% AI a fail?
20% is where Turnitin starts showing a number, not where misconduct starts. The company's stated false-positive rate — under 1% — applies specifically 'for documents with over 20% of AI writing', so below that line it is not claiming that accuracy at all. A department treating 20% as a limit has adopted Turnitin's display threshold as a disciplinary one, which is not what it is.
What exactly is the percentage measuring?
Not how 'AI-like' your writing is. Turnitin defines it as 'the amount of qualifying text within the submission that Turnitin's AI writing detection model determines was likely generated by AI or likely generated and modified by an AI paraphraser or bypasser'. Qualifying text means prose sentences in long-form writing — lists, bullet points and non-sentence structures are excluded, so the denominator is not your whole document.
How do I read a GPTZero result?
GPTZero returns one of three document classifications — HUMAN_ONLY, MIXED or AI_ONLY — alongside a probability it describes as 'the chance that the detector is correct in its classification'. That is a confidence figure about the tool, not a proportion of your document. It also states that where its confidence category is high, '99.1% of human articles are classified as human, and 98.4% of AI articles are classified as AI' — figures from its own testing.
Two tools gave me different scores. Which is right?
Possibly neither, and the disagreement is expected rather than a malfunction. The tools measure different things on different scales — a Turnitin percentage of qualifying text is not comparable with a GPTZero confidence probability — and peer-reviewed testing found accuracy varying widely by tool and by how the text was produced. Two numbers that disagree are two weak signals, not one strong one.

Sources

  1. 1.Turnitin's AI writing detection capabilities FAQs Turnitin Guides, 2026
  2. 2.GPTZero FAQ — classifications, confidence, and how results should be used GPTZero, 2026
  3. 3.Understanding false positives within our AI writing detection capabilities Turnitin, 2023
  4. 4.Testing of detection tools for AI-generated text Weber-Wulff et al. — International Journal for Educational Integrity 19:26, 2023