humanflow
AI detection · The HumanFlow team · 16 min read

Can you trust a 98% accuracy claim? How to read AI-detector marketing

"98% accurate" usually means "98% under conditions in the footnote." What accuracy claims really measure, the base-rate math, and a checklist for any vendor page.

Not at face value — and not because vendors are lying. Nearly every published accuracy figure is true under conditions stated somewhere in the fine print: a minimum AI percentage, unedited text, older models, English only, or the vendor's own test set. This page teaches you to find those conditions, do the base-rate arithmetic yourself, and evaluate any detector's number in about five minutes.

Every claim quoted below comes from a vendor page we fetched while writing this, with its conditions shown. Nothing is paraphrased into something scarier or safer than the original. That's the whole method, and by the end you'll be able to run it on any detector — including ours, which is why we publish no accuracy percentage without published methodology. The same rule you should hold vendors to, we hold ourselves to.

Three numbers hide inside every "accuracy" claim

A detector makes a yes/no call, and reality is yes/no, so every scan lands in one of four boxes: AI text flagged (true positive), AI text missed (false negative), human text passed (true negative), human text flagged (false positive). Every honest statistic about a detector is some ratio of those four boxes. Marketing usually gives you one ratio and lets you assume it describes all of them.

Sensitivity (true-positive rate): of the AI texts, what share got caught? This is the number vendors love, and on pristine text the best tools do earn something like it: Pangram was the only one of four to hold a false-positive rate at or below 0.5% in Jabarian and Imas's NBER working paper, and GPTZero recorded 2.50% false negatives in the DUPE study. Note what that sentence does not say. It names two tools, in two papers, neither peer-reviewed. It is not a rate for the category, and the largest peer-reviewed comparison to date put all fourteen tools it tested below 80%. High sensitivity is real, worth having, and unevenly distributed.

Specificity (true-negative rate): of the human texts, what share passed correctly? Its complement is the false-positive rate — the probability an innocent writer gets flagged. This is the number that determines whether the tool is safe to point at people, and it's the one most often missing, footnoted, or measured under favorable conditions.

Accuracy: correct calls divided by all calls. It sounds like the master number. It's the weakest of the three, because it depends entirely on the mix of the test set. Feed a detector 5,000 AI texts and 5,000 human texts and accuracy is the average of sensitivity and specificity. Feed it 9,500 AI texts and 500 human texts, and a detector that flags everything — every human included — scores 95% "accuracy." Statisticians call this the accuracy paradox. Any vendor page that says "accuracy" without telling you the test-set mix has told you almost nothing.

There's a fourth number nobody advertises, and it's the one that actually answers your question. Precision: of the texts that got flagged, what share were really AI? That depends not just on the detector but on how much AI text is in the pile you scan — the base rate. We'll compute it below, because it's the difference between "99% accurate" and "one flag in three is wrong," and both can be true of the same tool at the same time.

The conditional fine print, in the vendors' own words

Here is the trick — though "trick" is unfair to the vendors who state their conditions plainly. The pattern: a big number in the headline, and the conditions that scope it somewhere less prominent. Reading the two together is a skill. Practice on these real examples, quoted exactly.

Turnitin: 98% — for documents already heavily flagged

Turnitin's flagship claims are 98% accuracy and a sub-1% false-positive rate. Its AI-detector page states the condition in the same sentence, to its credit: "Our AI writing detector has a false positive rate of less than 1% for documents containing more than 20% AI-generated content." Its support documentation goes further: "We strive to maximize the effectiveness of our detector while keeping our false positive rate... under 1% for documents with over 20% of AI writing" (Turnitin, fetched for this article).

Read the condition twice. The false-positive guarantee applies only to documents the detector itself has already flagged above 20%. Below that line, Turnitin publishes no error rate at all — and its interface concedes why: "no score or highlights are attributed for AI detection scores in the 1% to 19% range. When AI is detected below the 20% threshold in the report, it is now indicated with an asterisk (*%) and no percentage is attributed" (Turnitin support documentation). The asterisk is Turnitin admitting, in the product itself, that low-range scores aren't reliable enough to print. That's responsible engineering — genuinely, and better than showing a shaky number as if it were solid. But it means the 98% is a claim about the detector's most confident cases, not about your document.

The scope conditions continue: at least 300 words of long-form prose, under 30,000 words, and — a detail almost nobody quotes — "Turnitin's AI writing detection capabilities are able to detect likely AI-generated content for documents submitted in long-form English, long-form Spanish and long-form Japanese. However, our AI paraphrasing and bypassing detection capabilities are only available for English submissions" (Turnitin support documentation). Three languages for detection; one for paraphrase detection. The headline number said none of this.

GPTZero: 99% — with the qualifier doing heavy lifting

GPTZero's homepage leads with "99% Accuracy" and elaborates: a "99% accuracy rate when spotting AI-generated text vs. human writing." But GPTZero also publishes the underlying benchmark result, and it's more interesting than the headline: "detecting 95.7% of AI texts while only incorrectly predicting 1% of human texts as AI" on the RAID benchmark — with "accuracy that jumps to over 99% when filtering to modern LLMs like GPT4" (gptzero.me, fetched for this article).

Watch the conditional: 99% arrives when filtering to modern LLMs — that is, after excluding the generators the detector handles worst. The unfiltered figure is 95.7% sensitivity at 1% false positives. That is a strong result, and naming RAID puts GPTZero ahead of rivals who cite nothing at all. Apply this article's own test to it, though, because the distinction is exactly the one this piece exists to teach: RAID is an independent, peer-reviewed benchmark with a public test set, and 95.7% is GPTZero's report of its own run against it. The figure does not appear in the paper. That is not an accusation — self-reporting against a public benchmark is far better than self-reporting against a private one, since anyone can rerun it. It is simply not the same thing as an independent measurement, and a vendor citing a benchmark is still a vendor citing itself. GPTZero also publishes a separate 96.5% figure for mixed human-AI documents, and this warning: "No AI detector is 100% accurate, and AI itself is changing constantly. Results should not be used to punish or as the final verdict" (gptzero.me). The vendor's own footer refutes the way many customers use the product.

Originality.ai: 99% — measured by Originality.ai

Originality.ai claims 99% accuracy on the latest models and calls itself "the most accurate AI Detector" citing "multiple independent studies." Its accuracy page publishes per-model detail: the Lite model at 98.91% true positives with 0.52% false positives; Turbo at 99.11% with 1.1% — measured on internal benchmarks of up to 456,872 samples (Originality.ai accuracy study, fetched for this article). The condition here isn't a threshold; it's the authorship of the evidence. The company built the test sets, ran the tests, and published the results — including its frequent head-to-head studies of competitors. Transparent, detailed, and still self-graded homework.

And yet the same page carries the most quotable caveat in the industry: "Don't trust a SINGLE 'accuracy' number without additional context," and the admission that false-positive rates are "still too high to be relied upon for disciplinary action." Full analysis in our Originality.ai review.

Winston AI: 99.98% — the best-column number

Winston AI's "99.98% accuracy rate" is the biggest headline figure in the market. Its own transparency post shows what it measures: in a self-built test of 5,000 AI texts and 5,000 human texts, Winston caught 99.98% of the AI texts. Its accuracy on the human texts was 99.50% — a 0.5% false-positive rate, about 25 innocent texts per 5,000. The AI samples came from GPT-3.5/4-era and Claude v1/v2-era models; the human samples were pre-2021; every text was at least 600 characters (Winston AI transparency post, December 2023). The headline quotes the best cell of the confusion matrix, from a clean test, on models that are now several generations old. Details in our Winston AI review.

The pattern, condensed

VendorHeadlineThe condition, in their own words or oursWhat's not covered
Turnitin98% accuracy, <1% false positives"...for documents containing more than 20% AI-generated content"; 300+ words; English/Spanish/Japanese long-formThe 1–19% range (asterisked as unreliable); short texts; paraphrase detection beyond English
GPTZero"99% Accuracy"95.7% detection at 1% FP, self-reported against RAID; 99% "when filtering to modern LLMs like GPT4"The models filtered out; the vendor's own "not the final verdict" warning
Originality.ai99% on latest models98.91–99.11% TP / 0.52–1.1% FP on internal benchmarks it built itselfIndependent replication; its own page says don't trust one number
Winston AI99.98% accuracyAI-side rate on a self-built 10,000-text set; human-side was 99.50%; models were GPT-3.5/4, Claude v1/v2; pre-2021 human textCurrent models; edited or paraphrased text; anything independently verified

Four vendors, four different kinds of condition: a score threshold, a model filter, self-authored evidence, and a best-column selection. All four numbers are, as scoped, plausibly true. None of them is the number you probably thought you were reading.

The base-rate walkthrough: do this arithmetic before trusting any flag

Here is the part that no vendor page will compute for you, using no source but arithmetic. Suppose a detector is genuinely excellent: 95% sensitivity (catches 95 of every 100 AI texts) and 99% specificity (a 1% false-positive rate). Take those as given — no skepticism, vendor's best case.

Now scan 10,000 documents where 10% are actually AI-written — roughly the prevalence Turnitin reported for meaningful AI presence in its first year of real submissions.

  • 1,000 documents are AI. The detector catches 95% of them: 950 true flags. It misses 50.
  • 9,000 documents are human. At a 1% false-positive rate: 90 innocent writers flagged.
  • Total flags: 1,040. Of those, 90 are wrong.

So 90 ÷ 1,040 ≈ 8.7% of all flags land on someone who did nothing — roughly one false accusation for every ten or eleven true catches, from a detector performing exactly as advertised. Not one error per ten thousand. One per eleven flags.

Now drop the prevalence. Scan a pool where only 1% is AI — a trusted writing team, an honors seminar, any low-cheating population:

  • 100 AI documents → 95 caught.
  • 9,900 human documents × 1% → 99 falsely flagged.
  • Flags: 194. Wrong flags: 99. A slight majority of accusations are false.

Same detector. Same honest 99%-specific spec. The flag went from "probably right" to "coin flip" because of nothing but the population it was pointed at. This is why a detector can be simultaneously "99% accurate" and wrong about half the people it accuses; the accuracy claim describes the boxes, and precision — the thing you actually care about when a flag appears — depends on the base rate the vendor cannot know. It's the same math that led Vanderbilt to disable Turnitin's indicator in August 2023 after running the false-positive arithmetic at institutional scale, and it's developed in full in our false positives explainer.

One more turn of the screw: false positives don't distribute randomly. The documented risk order runs non-native English speakers first (Liang et al. in Patterns, 2023, found seven detectors falsely flagged an average of 61.22% of human-written TOEFL essays while judging native-speaker 8th-grade essays near-perfectly), then writers trained into rigid structures, technical writers, and heavy self-editors. Your 90 innocent flags aren't spread evenly across the 9,000; they cluster on predictable people. A vendor's aggregate false-positive rate can be honest and still conceal a subgroup rate ten times higher — which is exactly what the Patterns study measured.

Why the gap between benchmark and reality keeps reopening

Two structural reasons, worth knowing so the next impressive number doesn't reset your skepticism.

First, thresholds are business decisions. Every detector converts a continuous statistical signal — how machine-typical the text is, measured through predictability and structural variation — into a verdict by drawing a line. Draw it low, you catch more AI and flag more humans; draw it high, the reverse. The line's position is a product-management choice about which error the vendor would rather explain, which is why identical text scores differently across tools (see how detectors work). An accuracy claim is downstream of a threshold choice you can't see and the vendor can move.

Second, benchmarks age instantly and adversaries don't sit still. The RAID benchmark (Dugan et al., ACL 2024) tested 12 detectors against 6 million-plus generations across 11 models, 8 domains, and 11 adversarial attacks, and concluded that "current detectors are easily fooled by adversarial attacks, variations in sampling strategies, repetition penalties, and unseen generative models" — while noting pointedly that commercial detectors "claim to detect machine-generated text with extremely high accuracy (99% or more)." A clean-benchmark number is a ceiling, measured on yesterday's models and nobody trying to evade. The canonical cautionary tale remains OpenAI's own classifier: 26% of AI text correctly identified, 9% of human text falsely flagged, retired by OpenAI in July 2023 for low accuracy. The company that built the generator couldn't reliably detect it — a data point that should calibrate how much confidence any third party's decimal places deserve.

To be fair to the whole category: detection is not snake oil. Raw, unedited AI text really is caught most of the time, by most serious tools. The dishonesty isn't in building detectors; it's in compressing "strong performance under narrow conditions" into a context-free percentage — and in customers using that percentage as a verdict on a human being.

The checklist: eight questions for any vendor's number

Bookmark this section. When a detector page shows you a percentage, ask:

  1. Which ratio is it? Sensitivity, specificity, or blended "accuracy"? If the page won't say, assume it's the most flattering of the three.
  2. What's the false-positive rate, stated separately? No FP number anywhere is disqualifying for any use involving people. A 99% headline with a hidden 5% FP rate is a machine for wrongful accusations.
  3. What population, and what conditions? Threshold restrictions (Turnitin's ">20%"), model filters (GPTZero's "modern LLMs"), length minimums, language limits. The claim is only valid inside its stated envelope — check whether your text is inside it.
  4. Which generators, and how old? A figure measured on GPT-4-era output tells you little about this year's models. Look for the test date and model list; absence is an answer.
  5. Was the text edited? Nearly all published claims are for raw output. Human-edited, mixed, translated, and paraphrased text is where benchmarks show performance collapsing — and where most real text now lives.
  6. Who built the test set? Vendor-built benchmarks — including every number in Winston's and Originality.ai's headlines — are unaudited. Independent, adversarial, peer-reviewed evaluations (RAID is the current reference) outrank them.
  7. Does the arithmetic survive your base rate? Multiply the FP rate by your human volume; compare to expected true catches. Ten minutes with the walkthrough above tells you your real precision.
  8. Does the vendor's own fine print contradict the sales pitch? Turnitin asterisks low scores. GPTZero says results "should not be used to punish." Originality.ai says one number is never enough. When the vendor tells you the limits of its product, that's the most reliable sentence on the page. Believe it over the headline — and over anyone, including us, selling in this market.

A claim that survives all eight questions is rare. When you find one, you've found a vendor worth taking seriously — and you should still treat the score as one input beside process evidence, drafts, and conversation, never as the verdict.

FAQ

What does "98% accuracy" actually mean on a detector's site? By itself, almost nothing — accuracy depends on the test set's mix of AI and human text, the models used, and the conditions applied. Turnitin's 98%, for instance, is scoped to documents already flagged above 20% AI, in three supported languages, at 300+ words. Always find the ratio (sensitivity vs. specificity) and the test conditions before crediting the number.

Can a detector be 99% accurate and still wrong about most people it flags? Yes — that's the base-rate problem. If only 1% of scanned documents are AI, a 95%-sensitive, 99%-specific detector produces about 99 false flags for every 95 true catches: most accusations wrong, spec sheet fully honest. Precision depends on prevalence, which no vendor can put on a sales page.

What's the difference between sensitivity and specificity? Sensitivity is the share of AI texts correctly caught; specificity is the share of human texts correctly passed. A detector can have superb sensitivity and mediocre specificity — great at catching AI, dangerous to innocent writers. Marketing headlines are usually sensitivity; the risk to real people lives in specificity.

Are any AI detector accuracy claims independently verified? Some. GPTZero cites results on RAID, the peer-reviewed ACL 2024 benchmark, which is the field's main independent evaluation. Most other headline figures — Winston's 99.98%, Originality.ai's 99% — come from vendor-built internal tests. Independent adversarial benchmarks consistently show lower, more fragile performance than vendor tests.

Why did OpenAI shut down its own AI detector? OpenAI's classifier identified only 26% of AI-written text and falsely flagged 9% of human writing, and OpenAI retired it in July 2023, citing low accuracy. It remains the clearest evidence that detection is hard even with full knowledge of the generator — useful calibration against any 99%+ claim.

Do accuracy claims apply to text written in other languages? Usually not, or only partially. Turnitin supports long-form English, Spanish, and Japanese — with paraphrase detection in English only. Claims measured on English text routinely get quoted as if universal. Non-native English writing is also the best-documented false-positive risk (61.22% average false-flag rate on human TOEFL essays across seven detectors in Liang et al., 2023).

What should I use instead of trusting one score? Triangulate: run more than one detector and note disagreement, weigh process evidence (drafts, revision history, the writer's track record), and treat every score as a probability about text patterns — not authorship. Vendors themselves say this; GPTZero's page states results "should not be used to punish or as the final verdict."

Key facts

  • Turnitin's <1% false-positive claim is conditional: "for documents containing more than 20% AI-generated content"; scores of 1–19% display as an asterisk with "no percentage attributed" (Turnitin, fetched 2026).
  • Turnitin detection covers long-form English, Spanish, and Japanese only; paraphrase/bypass detection is English-only; 300-word minimum (Turnitin support documentation).
  • GPTZero's self-reported figures against the RAID benchmark: 95.7% of AI texts detected at 1% false positives, rising "to over 99% when filtering to modern LLMs like GPT4" (gptzero.me, fetched 2026).
  • Winston AI's 99.98% is its AI-side rate from a self-built 10,000-text test (December 2023) using GPT-3.5/4 and Claude v1/v2 outputs; the human-side rate was 99.50% — a 0.5% false-positive rate.
  • Originality.ai reports 98.91–99.11% true positives and 0.52–1.1% false positives on internal benchmarks of up to 456,872 samples — and warns "Don't trust a SINGLE 'accuracy' number without additional context."
  • OpenAI retired its own classifier in July 2023 at 26% detection and 9% false positives.
  • RAID (ACL 2024): 12 detectors, 6M+ generations, 11 adversarial attacks — detectors "easily fooled" by paraphrase, sampling changes, and unseen models despite 99%+ marketing claims.
  • Base-rate arithmetic: at 95% sensitivity / 99% specificity and 10% AI prevalence, ~1 in 11 flags is a false accusation; at 1% prevalence, most flags are.

Sources

  1. Turnitin — AI detector page and AI-writing support documentation (false-positive condition, asterisk display, word minimums, language support), fetched for this article.
  2. GPTZero — gptzero.me homepage (accuracy claims, RAID citation, limitations statement), fetched for this article.
  3. Originality.ai — accuracy study page and homepage (per-model benchmark figures, caveats), fetched for this article.
  4. Winston AI — "Setting new standards in AI content detection," December 5, 2023 (dataset methodology behind the 99.98% claim).
  5. Dugan et al. — "RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors," ACL 2024 (arXiv:2405.07940).
  6. OpenAI — AI text classifier retirement announcement, July 2023.
  7. Liang et al. — "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023.
  8. Vanderbilt University — statement on disabling Turnitin's AI detection, August 2023.
All postsPublished by The HumanFlow team