humanflow

Do AI detectors work on the newest models?

Unevenly — one detector performed very well on medium-to-long passages in the most recent systematic comparison, while others were placed in a tier the authors called unsuitable for short text and susceptible to humanizers.

Last reviewed 15 August 2026 · The HumanFlow team

The most current large test is an NBER working paper by Jabarian and Imas, which ran 1,992 human and 1,992 AI passages across four frontier models, plus text put through a commercial humanizer. It is a working paper rather than peer-reviewed work, and it says so on the cover — worth stating plainly, because it is quoted widely without that qualifier.

Its headline finding is a split rather than a verdict on the field. It reports Pangram reaching "essentially zero FPRs and FNRs" on medium-to-long passages, while placing GPTZero and Originality.ai in a "secondary tier" that is unsuitable for very short text and "susceptible to 'humanizers.'" So "do detectors work on new models" has no single answer: it depends heavily on which detector, and on how long the passage is.

Two conditions degrade every result in that study. Short text is the first — the shorter the passage, the less signal there is to measure, and the secondary tier fails there specifically. Paraphrasing is the second, and it is the finding that recurs across the whole literature: Weber-Wulff and colleagues measured 26% accuracy across fourteen detectors on machine-paraphrased text.

The paper's framing of the field is the part most worth carrying away: "published claims about detector accuracy are hard to verify: they are often based on private data, focus on a single threshold." That applies to vendor claims about new models as much as to older ones, and it is the reason a detector's marketing page is not evidence that it handles the model you are asking about.

When this answer changes

It changes every time either side ships. A new model release can move results without any detector changing, and detector updates arrive without changelogs, so a result describes a moment rather than a property.

It changes most for short text. If your question is about a paragraph rather than an essay, the honest answer is weaker than anything above — that is the specific condition the secondary tier was found unsuitable for.

Where to go next

Sources

One of our direct answers.