humanflow
AI detection · The HumanFlow team · 16 min read

Are AI Detectors Reliable in 2026? What Every Published Study Actually Shows

No peer-reviewed study supports the accuracy vendors imply — the largest put all 14 tools under 80%. Every published reliability finding since 2023, reviewed.

AI detectors in 2026 are least bad at exactly one job: flagging long, unedited, English-language AI text. Even there the evidence is thinner than the marketing — the largest peer-reviewed comparison tested fourteen tools and found every one of them below 80% accuracy, with only five above 70%. Outside that narrow lane — edited drafts, non-native English writers, short passages — the published record gets considerably worse. Reliable enough to inform a conversation; not reliable enough to end one.

That's the summary. This article is the evidence: every significant published reliability finding since 2023, assembled in one place, including the numbers vendors would rather you read without their conditions and the numbers critics quote without their context. Both sins are common, and the field can't be understood while committing either.

The published record, in one table

FindingNumberConditions attachedSource, year
OpenAI's own classifier: AI text caught26%Its maker retired it for low accuracyOpenAI, July 2023
OpenAI's classifier: human text falsely flagged9%Same tool, same announcementOpenAI, July 2023
Human TOEFL essays falsely flagged (avg. of 7 detectors)61.22%Non-native English writers; 91 essaysLiang et al., Patterns, 2023
TOEFL essays flagged by at least one of the 789 of 91Same studyLiang et al., 2023
TOEFL essays flagged by ALL seven detectors18 of 91Same studyLiang et al., 2023
Same detectors on US 8th-grader essaysNear-perfectNative-speaker text, same toolsLiang et al., 2023
Turnitin claimed accuracy98%Only where >20% of the document is flaggedTurnitin FAQ
Turnitin claimed false positive rate<1%Same >20% condition; 1–19% shows as asteriskTurnitin FAQ
AI text Turnitin deliberately leaves unflagged~15%"We find about 85% of it. We let probably 15% go by"Annie Chechitelli, Turnitin CPO, BestColleges, Apr 2023
Every one of 14 detectors testedunder 80%"only 5 over 70%"; 54 cases, 756 testsWeber-Wulff et al., IJEI 19:26, 2023
Same 14 tools on machine-paraphrased AI text26%Paraphrased output, all tools pooledWeber-Wulff et al., 2023
7 detectors, baseline accuracy39.5%Before any adversarial techniquePerkins et al., IJETHE 21:53, 2024
Same 7 after adversarial editing22.2%17.4 is the mean reduction, not the accuracyPerkins et al., 2024
Detectors meeting a ≤0.5% false-positive bar1 of 4 (Pangram)Not peer-reviewed — NBER working paperJabarian & Imas, NBER 34223, 2025
Papers ≥20% AI in Turnitin's first year~11% of 200M+Self-reported vendor telemetryTurnitin, April 2024

Every number below gets unpacked with its context. But notice the shape of the table before the details: the good numbers all carry conditions about unedited text and native English, and the bad numbers all live exactly where those conditions fail. That's not a coincidence. It's the whole story.

What "reliable" would even mean

Before judging the tools, define the test. A detector you could honestly call reliable would need five properties at once, and it helps to name them because vendors quote whichever one flatters them:

Sensitivity — catches AI text when present. Specificity — doesn't flag human text. Calibration — a "73% AI" score means something consistent; 73 isn't just a vibe with decimals. Robustness — holds up when text is edited, paraphrased, translated, or generated by a model released after the detector's last training run. Equity — error rates don't concentrate on identifiable groups of writers.

The 2023–2026 record shows real progress on sensitivity, genuine engineering effort on specificity, and persistent, documented failure on robustness and equity. Calibration remains mostly unexamined — no major vendor publishes calibration curves at all. Hold that five-part frame; every study below scores against it.

2023: the annus horribilis the field still hasn't outrun

Three findings from 2023 define the floor of this debate, and all three still get cited because nothing has formally superseded them.

OpenAI retired its own detector. In January 2023 OpenAI released an AI-text classifier; by July 2023 it had quietly killed it, citing low accuracy. The published numbers were brutal: it correctly identified only 26% of AI-written text while falsely flagging 9% of human writing. This is the canonical fact of the field — the company with the deepest possible knowledge of how its models generate text could not build a reliable detector of that text and said so. Every vendor claiming 99% today is claiming to have solved, from the outside, a problem the model's own maker abandoned from the inside. Maybe some have made real progress. The burden of proof sits with them.

Liang et al. documented systematic bias. Published in Patterns (Cell Press, 2023), the Stanford-led study ran 91 human-written TOEFL essays through seven detectors. The average false-positive rate was 61.22%. Eighty-nine of the 91 essays were flagged as AI by at least one detector; 18 were flagged by all seven — meaning a fifth of these real students would have been "caught" no matter which tool their institution happened to license. The same seven detectors were near-perfect on essays by native-speaking US 8th graders. The mechanism is mundane and therefore damning: non-native writers use narrower vocabulary and safer syntax, which reads as low perplexity — the same statistical signature as machine text. (Turnitin was not among the seven tested; it uses the same statistical family, which is why the study appears in Turnitin debates — with that caveat, always.)

The Washington Post caught a false positive in a live demo. In April 2023, Geoffrey Fowler tested Turnitin's then-new indicator and it flagged an innocent student's genuine writing, while also struggling with mixed human/AI drafts. Turnitin acknowledged blended documents are the hard case. That acknowledgment aged well — mixed text is still the hard case, three years later.

The institutional response followed the evidence. Vanderbilt University disabled Turnitin's AI indicator in August 2023 and published its reasoning, which was arithmetic rather than ideology: even at Turnitin's claimed sub-1% false positive rate, screening tens of thousands of papers implies hundreds of innocent flags a year, each a potential misconduct case built on a probability score. Several other large universities did the same or demoted the score to advisory.

The base-rate problem: why "99% accurate" still accuses the innocent

This section is arithmetic, not opinion, so no study is needed — you can check it on paper.

Take a detector that catches 95% of AI text (sensitivity) and correctly clears 99% of human text (specificity) — better than most published evidence supports, but grant it. Run it on 10,000 student papers in a course where 10% of submissions actually contain AI text.

  • 1,000 papers are AI-assisted → the detector flags 950 of them. Good.
  • 9,000 papers are fully human → 1% get flagged anyway → 90 innocent students.

So 90 of the 1,040 flags — roughly one in eleven — point at students who did nothing. Every tenth accusation lands on an innocent person, using a detector better than the published evidence says exists, and that ratio worsens as honest students become the majority: at a 5% AI rate, nearly one flag in six is false. This is why "the false positive rate is under 1%" and "false accusations will be routine" are both true at once, and why Vanderbilt's math holds regardless of which vendor's brochure you prefer. Any policy that treats a flag as a verdict has decided, in advance, to sacrifice those students. The full walkthrough with different assumptions lives on our false positives page.

What actually improved since 2023

Honesty cuts both ways, and the field's critics often recycle 2023 numbers as if nothing has happened since. Real things changed.

The best tools got genuinely good at raw AI text — and "the best tools" is doing real work in that sentence. Two recent evaluations found near-ceiling performance on unedited output, but each for a single tool: Pangram was the only one of four to meet a false-positive rate at or below 0.5% in Jabarian and Imas's NBER working paper, and GPTZero recorded a 0.00% false-positive rate with 2.50% false negatives in Weichert and Dimobi's ZeroGPT study. Neither paper is peer-reviewed, and neither generalizes to the category — in the same DUPE test ZeroGPT managed 78.62% overall with a 24.64% false-positive rate. If someone pastes an unmodified chatbot answer into a submission, a good detector will usually catch it. That is a claim about a few named tools on pristine text, not about detectors.

Vendors engineered against false positives — some of them, anyway. Turnitin's design deserves specific credit: its 98%/<1% claims are explicitly scoped to documents over the 20% threshold; scores of 1–19% display as an asterisk rather than a number, an open admission that low-range scores are noise; it refuses to score fewer than ~300 words; and its chief product officer, Annie Chechitelli, described the trade-off to BestColleges in April 2023 as "we find about 85% of it. We let probably 15% go by". That is a vendor pricing in its own fallibility — responsible engineering, and we say so as a competitor.

Benchmarks arrived. The RAID benchmark (ACL 2024) gave the field its first serious public adversarial evaluation — millions of documents spanning models, domains, decoding strategies, and eleven evasion attacks. Vendors now compete on it, which beats competing on adjectives. The catch: QuillBot and Grammarly both cite RAID for ~99% claims, and Grammarly claims first place on it — the same benchmark yielding multiple winners tells you how much slicing the claims involve.

Watermarking became real. Google DeepMind's SynthID text watermarking opened to developers in 2024 and runs in Gemini output. It's a structurally different technology — the generator embeds a statistical signature at creation, so checking is closer to proof than inference. Its limits are equally structural: it requires the model vendor's cooperation, does nothing for models that don't watermark, and degrades under heavy paraphrase. It hasn't displaced statistical detection, but it's the first genuinely new reliability idea since 2023.

Scale data emerged. Turnitin's first-anniversary release (April 2024) reported 200M+ papers screened, ~11% with ≥20% AI writing and ~3% at 80%+. Vendor telemetry, not neutral research — but it's the largest dataset in existence on real-world AI-writing prevalence, and it anchors the base-rate math above in observed reality. Turnitin has published later prevalence updates, though not on a consistent enough basis to build a trend line from.

What hasn't improved: the arms race nobody is winning

Now the other ledger.

Edited and mixed text is still the unsolved case. The Washington Post identified blended drafts as the weakness in April 2023; in 2026, Grammarly's own detector page still concedes that "lightly edited AI-generated text may evade detection." Every good number in this article belongs to unedited text. But almost nobody submits unedited text anymore — real documents are drafted, expanded, trimmed, and rephrased by humans and machines in alternating passes. Detection research keeps testing the case that's disappearing from the world.

The humanizer-vs-detector loop is a treadmill. An entire industry — ours included, disclosed plainly — sells paraphrasing and rewriting tools, and detectors retrain against their output; Turnitin added paraphrase detection in July 2024 and bypasser detection in August 2025, English-only, with no accuracy figure published for either. Then rewriting tools adjust, then detectors retrain. RAID formalized the finding: adversarial paraphrase degrades every statistical detector tested, by amounts that vary but never reach zero. There is no equilibrium coming, because both sides are optimizing against the other's last move. This is why our own humanizer doesn't promise to beat any detector — nobody can honestly promise that, on either side of the treadmill, and the vendors who do are selling a snapshot of a moving target.

The equity problem has not been re-measured at scale. Liang et al. is three years old, and detectors have retrained since — but no comparably rigorous published replication on non-native writers exists. Vendors claim improvement; the study that would demonstrate it hasn't been published. The documented risk order still stands: non-native English speakers, students taught rigid essay structures, technical and legal writers, heavy self-editors, and neurodivergent writers whose prose reads as statistically regular. Absence of new evidence is not evidence of repair.

New models keep resetting the clock. Every frontier release changes the statistical texture detectors were trained on. GPTZero's own claim structure shows the mechanism: "over 99% when filtering to modern LLMs" — accuracy is a function of which models you test, measured after detectors catch up. Reliability against last year's models is the only kind that can be measured; reliability against next month's is a promise.

Calibration and thresholds remain vendor secrets. Every score you see is a probability pushed through a threshold the vendor chose — a business decision balancing embarrassing false positives against embarrassing misses. That's why the same essay scores 12% on one tool and 74% on another, and why cross-tool disagreement is structural, not a bug being fixed. Mechanics in how detectors work.

The verdict, by use case

"Are AI detectors reliable in 2026?" dissolves into better questions the moment you ask reliable for what.

For teachers deciding whether to open a conversation: yes, cautiously. A high score on a long document is a reasonable trigger for a non-accusatory chat and a look at draft history. Turnitin's own guidance says the score shouldn't be the sole basis for action; that's the correct standard for every tool.

For anyone making accusations, terminations, or misconduct findings: no. The base-rate arithmetic makes score-only decisions statistically guaranteed to hit innocent people, and the equity record shows those people won't be randomly chosen. Process evidence — version history, drafts, discussion of the work — outranks any percentage.

For students self-checking: partially. A scan shows whether your writing reads as machine-typical to one classifier — useful for spotting false-positive risk, useless for predicting your institution's tool, which runs different thresholds you can't see. (Free options ranked in our free-detector guide.)

For publishers screening at volume: yes as triage, no as judge. Detectors sort a pile efficiently; every consequential decision still needs a human reading and a writer's account of process. Tool choices in our best-detector framework.

For everyone: a detector score is one noisy instrument's opinion about statistical texture. In 2026 that opinion is worth hearing on long, unedited, English text, and worth very little elsewhere. The tools improved. The verdict "informative, never sufficient" hasn't moved since 2023 — and on the current evidence, it shouldn't.

Our own position, stated once: we build a detector, we publish no accuracy percentage for it without published methodology, and we'd apply this article's standards to ourselves before asking you to. The full evidence file lives at our accuracy hub.

FAQ

Are AI detectors reliable in 2026? Within a narrow lane, the best tools are: long, unedited, English AI text is caught reliably by the strongest performers, though the largest peer-reviewed comparison still put all fourteen tools it tested below 80%. Outside that lane, no: edited and mixed text often slips through, and false positives concentrate hard on non-native English writing — 61.22% of human TOEFL essays in Liang et al. Reliable enough to prompt questions, not to settle them.

Have AI detectors improved since 2023? Substantially, at the top of the market — from OpenAI's retired classifier catching 26% of AI text to purpose-built tools now posting near-ceiling numbers on unedited output in the NBER and DUPE evaluations. Less so on the hard problems: mixed human/AI drafts, paraphrased text, and bias against non-native writers, where the 2023 findings have not been convincingly re-measured, let alone overturned.

How often are AI detectors wrong? Depends entirely on the text. On unedited AI output: wrong maybe 5–10% of the time. On human writing by non-native English speakers, the landmark study found detectors wrong 61.22% of the time on average. On lightly edited AI text, vendors themselves concede detection may fail. There is no single error rate — that's the point.

Can AI detectors be used as proof of cheating? No. Turnitin itself says its score shouldn't be the sole basis for an accusation; the base-rate math shows even a 99%-specific detector produces roughly one false flag per ten true ones under realistic conditions. Detector output can justify a conversation and a request for drafts — not a finding.

Why did OpenAI shut down its own AI detector? Low accuracy, by its own July 2023 announcement: 26% of AI text identified, 9% of human text falsely flagged. It remains the field's defining fact — the model's own maker couldn't reliably detect its output from the outside either.

Do AI detectors discriminate against non-native English speakers? The published evidence says the risk is real and structural: Liang et al. (Patterns, 2023) found seven detectors falsely flagged an average of 61.22% of human-written TOEFL essays while scoring near-perfectly on native 8th-graders' essays. Simpler vocabulary and safer syntax mimic the statistical signature of AI text. No comparably rigorous replication has shown the problem fixed.

Will watermarking make AI detection reliable? It could change the game for cooperating models — Google's SynthID embeds a checkable signature in Gemini output, which is closer to proof than statistical guessing. But it requires the generator's cooperation, covers nothing generated by non-watermarking models, and degrades under heavy paraphrase. It complements statistical detection; it doesn't replace it yet.

Can humanizers reliably beat detectors? No — and detectors can't reliably beat humanizers either. RAID showed adversarial paraphrase degrades every detector tested; detectors then retrain, and rewriting tools adjust again. It's a treadmill with no finish line, which is why we refuse to publish "bypass rates" for our own humanizer: any such number describes one moment of a moving fight.

Key facts

  • OpenAI's own AI text classifier caught only 26% of AI writing and falsely flagged 9% of human writing; OpenAI retired it in July 2023 citing low accuracy (OpenAI).
  • Liang et al. (Patterns, 2023): seven detectors falsely flagged an average of 61.22% of 91 human TOEFL essays; 89/91 flagged by at least one tool, 18/91 by all seven; near-perfect on native US 8th-grader essays.
  • Turnitin's 98% accuracy and <1% false-positive claims apply only above its 20%-AI threshold; 1–19% displays as an asterisk; ~300-word minimum (Turnitin FAQ).
  • Turnitin screened 200M+ papers April 2023–April 2024: ~11% had ≥20% AI writing, ~3% were 80%+ AI (Turnitin, April 2024).
  • Vanderbilt University disabled Turnitin's AI indicator in August 2023, publishing false-positive arithmetic as the reason.
  • Weber-Wulff et al. (IJEI 19:26, 2023): all 14 tools tested scored below 80% accuracy, only 5 above 70%, and 26% on machine-paraphrased AI text.
  • Perkins et al. (IJETHE 21:53, 2024): 39.5% baseline accuracy across 7 detectors, falling to 22.2% under adversarial editing.
  • At 95% sensitivity and 99% specificity with 10% AI prevalence, roughly 1 in 11 flags is a false accusation — arithmetic, not estimate.
  • RAID (ACL 2024), the main public adversarial benchmark, found paraphrase attacks degrade every statistical detector tested.

Sources

  1. OpenAI, "New AI classifier for indicating AI-written text" — including the July 2023 retirement note and 26%/9% figures.
  2. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J. "GPT detectors are biased against non-native English writers." Patterns (Cell Press), 2023.
  3. Turnitin AI writing detection FAQ and transparency pages — accuracy conditions, asterisk policy, minimum length (turnitin.com).
  4. Turnitin first-anniversary press release, April 2024 — 200M+ papers reviewed, prevalence figures.
  5. BestColleges — interview with Annie Chechitelli, Turnitin chief product officer, April 2023.
  6. Fowler, G. "We tested a new ChatGPT-detector for teachers. It flagged an innocent student." Washington Post, April 3, 2023.
  7. Vanderbilt University Brightspace announcement, "Guidance on AI detection and why we're disabling Turnitin's AI detector," August 2023.
  8. Dugan, L., et al. "RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors." ACL 2024. Dugan, Hwang, Trhlík, Ludan, Zhu, Xu, Ippolito and Callison-Burch. arXiv:2405.07940.
  9. Google DeepMind, SynthID text watermarking announcements and developer release, 2024.
  10. Grammarly AI Detector page — "lightly edited AI-generated text may evade detection" and standalone-use warning (grammarly.com/ai-detector, fetched August 2026).
  11. GPTZero homepage — claim structure including "when filtering to modern LLMs" (gptzero.me, fetched August 2026).
  12. Jabarian, B. & Imas, A. (2025). Automated Detection of AI-Generated Text. NBER working paper 34223 — not peer-reviewed.
  13. Weichert, J. & Dimobi, D. (2024). DUPE: Detection Undermining via Prompt Engineering. arXiv:2404.11408 — preprint, originally a Virginia Tech CS 5914 course project; not peer-reviewed.
All postsPublished by The HumanFlow team