humanflow

AI detector accuracy: the research, reviewed

Six studies, tested between them on fourteen, eight, seven and four detectors, agree on one thing and disagree on almost everything else. They agree that accuracy collapses on paraphrased text. They disagree on how often human writing gets flagged, because they tested different tools on different corpora in different years.

Last reviewed 16 August 2026 · The HumanFlow team

The studies

StudyVenueScopePeer-reviewed
Liang et al. (2023)Patterns (Cell Press)7 detectorsYes
Weber-Wulff et al. (2023)Int. J. for Educational Integrity14 detectorsYes
Perkins et al. (2024)Int. J. of Educational Technology in Higher Education7 detectorsYes
Orenstrakh et al. (2023)arXiv preprint8 detectorsNo — preprint or working paper
Krishna et al. (2023)NeurIPS 36paraphrasing attackYes
Jabarian & Imas (2025)NBER Working Paper 342234 detectorsNo — preprint or working paper

What each one found

Liang et al. (2023)7 detectors
Flagged 61.22% of TOEFL essays written by non-native English speakers as AI-generated — 89 of 91 essays tripped at least one tool — while classifying essays by native-speaking US 8th-graders almost perfectly.
Weber-Wulff et al. (2023)14 detectors
Concluded the tools “are neither accurate nor reliable”. Every detector scored below 80% accuracy, and on machine-paraphrased text overall accuracy fell to 26%.
Perkins et al. (2024)7 detectors
Under adversarial modification, detectors caught 39% of AI-generated cases. The study counted an accusation as false whenever a human-written control was flagged at all.
Orenstrakh et al. (2023)8 detectors
Tested detectors against human and AI-generated work in computing education, where the formulaic conventions of the genre make the false-positive problem sharper.
Krishna et al. (2023)paraphrasing attack
Showed that paraphrasing AI output evades detectors — and, in the same paper, that retrieval over a store of previously generated text remains an effective defence against exactly that attack. The second half is quoted far less often than the first, including by people selling paraphrasers.
Jabarian & Imas (2025)4 detectors
Tested 1,992 human and 1,992 AI passages across four frontier models. Ranked Pangram first; placed GPTZero second, in a “secondary tier” that degrades on short passages and against humanizing tools.

The one finding they agree on

Detection accuracy collapses on paraphrased text. Weber-Wulff and colleagues measured 26% overall accuracy across all fourteen tools once text had been machine-paraphrased. Perkins and colleagues measured 39% under adversarial modification. Krishna and colleagues established the mechanism at NeurIPS before either — and proposed a defence against it in the same paper, which is the half of that title almost nobody quotes.

That convergence matters more than any single accuracy number, because it is the finding that survives the differences in method. It also has an uncomfortable implication for everyone in this market, including us: detection of rewritten text is an arms race, and a vendor selling a bypass rate is selling a snapshot of a moving target.

What the literature does not establish

That any specific tool has a specific accuracy today. Detector versions change without changelogs. A 2023 result describes software that no longer exists, and no study here re-tests on a schedule.

That the 61.22% figure applies to any one detector. It is a seven-detector average from Liang et al., and that study did not include Turnitin. The number is quoted against individual tools constantly, including against tools that were never in it. If you take one methodological point from this page, take that one.

That detectors are uniformly bad, or uniformly biased. Liang et al. found severe bias against non-native writers; Turnitin's own funded study reports near-identical rates for both groups above 300 words. Both findings are real, and we publish both on the accuracy page.

That any of this measures HumanFlow. None of these studies tested our detector or our humanizer. We are in the market they describe, not in the research.

How this page is maintained

Every finding above is stated as this site already states it elsewhere, with the same source attached. That is deliberate: summarising a paper from memory is how a miscitation enters a corpus, and once it looks cited it survives review. Where we could not verify an attribution, the claim is not on this page.

The question this page gets asked most, answered on its own: do AI detectors work on the newest models?

If a study is missing, or a finding here misrepresents one, tell us — corrections go in the public corrections log with the date and what changed.

If you are citing any of this, cite the papers rather than us — every one is linked above. How to cite this site explains why, and what we will not do to look more citable.

Common questions

What is the single best study on AI detector accuracy?
There is not one, and treating any of them as definitive is the error this page exists to prevent. Weber-Wulff et al. covers the most tools (fourteen); Liang et al. is the strongest evidence on bias against non-native speakers; Perkins et al. is the strongest on adversarial conditions. They tested different tools, different corpora and different versions, in different years.
Do any of these studies test the detector my institution uses?
Possibly not, and this is the most common misreading we see. The 61.22% figure comes from a seven-detector average in a study that did not include Turnitin, so it cannot be attributed to Turnitin — or to any single tool in it. Check which detectors a study actually tested before applying its number to yours.
Why do the accuracy figures vary so much between studies?
Because they measure different things. A study using unmodified AI output measures one world; a study using paraphrased or adversarially-modified text measures another. Detector versions also change without changelogs, so a 2023 result describes software that no longer exists.
Is a peer-reviewed study automatically more reliable here?
It is a meaningful signal and not a guarantee. We mark peer-review status on every row above because two of the most-cited results in this area are preprints or working papers — the NBER paper states its own non-peer-reviewed status — and they get quoted without that qualifier constantly.
Does HumanFlow appear in any of this research?
No. No independent study has tested our detector or our humanizer, and we are not going to imply otherwise by citing research that does not include us. That is also why we publish no bypass rate: we have no external measurement to publish, and an internal one about our own product would not be worth your trust.

Sources

  1. 1.GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press), 2023
  2. 2.Testing of detection tools for AI-generated text Weber-Wulff et al. — International Journal for Educational Integrity 19:26, 2023
  3. 3.Simple techniques to bypass GenAI text detectors: implications for inclusive education Perkins, Roe, Vu, Postma, Hickerson, McGaughran & Khuat — Int. J. of Educational Technology in Higher Education 21:53, 2024
  4. 4.Detecting LLM-Generated Text in Computing Education Orenstrakh, Karnalim, Suarez & Liut — arXiv:2307.07411, 2023
  5. 5.Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defense Krishna, Song, Karpinska, Wieting & Iyyer — NeurIPS 36, 2023
  6. 6.Artificial Writing and Automated Detection (NBER working paper, not peer-reviewed) Jabarian & Imas — NBER Working Paper 34223, 2025

Part of our research. See also the benchmark protocol we intend to run, and why we publish no bypass rates.