This page exists because the rest of this cluster argues that AI detection is unreliable, and honesty requires noting the case that cuts against it. Pangram is a detector rather than a humanizer, so it is not a commercial rival of ours — but it is a rival to our argument, which is a better reason to write carefully about it.
Pangram: what it measures, and what it proves
Pangram is an AI detector claiming 99.98% accuracy and a false positive rate of 1 in 10,000. Unusually for this category, an independent working paper broadly supports the direction of those claims on medium-to-long text. Also unusually, the company states its figures precisely enough to be argued with.
Last reviewed 15 August 2026 · The HumanFlow team
What the company claims
Read from pangram.com on 15 August 2026, the headline is “Detect AI-generated content with 99.98% accuracy”, with a more conservative “Pangram is over 99% accurate” in its FAQ. It attributes third-party verification to the University of Chicago and the University of Maryland, and publishes a comparison against GPTZero, Originality.ai and RoBERTa across six document types, attributed to research from Chicago Booth.
On false positives it is specific: the rate “at which human documents are incorrectly flagged as AI” is stated as “currently 1 in 10,000”. The model is described as trained on approximately one million human and AI documents.
Those are the vendor's claims about the vendor's product, and this site does not treat that kind of claim as established — a rule we apply to favourable numbers as well as unfavourable ones. What is different here is the precision. “1 in 10,000” is falsifiable in a way that “industry-leading accuracy” is not, and a company that publishes a checkable number has done something most of this category avoids.
What independent work found
The most current systematic comparison is an NBER working paper by Jabarian and Imas, which ran 1,992 human and 1,992 AI passages across four frontier models plus text put through a commercial humanizer. It reports Pangram reaching “essentially zero FPRs and FNRs” on medium-to-long passages, while placing GPTZero and Originality.ai in a “secondary tier” unsuitable for very short text and “susceptible to ‘humanizers.’”
That is meaningful corroboration and it carries two qualifiers. The paper is a working paper, not peer-reviewed, and states so itself. And its own framing of the field applies to Pangram as much as to anyone: “published claims about detector accuracy are hard to verify: they are often based on private data, focus on a single threshold.”
So the fair summary is that one tool has better independent evidence than the rest of a category whose evidence is generally poor. That is a real distinction and a modest one.
Known limitations
Short text. The documented advantage is on medium-to-long passages. Every detector loses accuracy as text shortens, because there is less to measure, and the tiering in the NBER paper was explicitly about length.
English. The evidence discussed above concerns English text. Nothing here establishes performance on other languages, and the best-known bias finding in this field — 61.22% of essays by non-native speakers flagged by seven detectors — is about writers rather than languages.
It is not what your institution runs. This is the practical limitation that matters most. If you have been flagged, it was almost certainly by Turnitin, GPTZero, Copyleaks or a similar tool, and Pangram's figures say nothing about the report in front of you.
What a Pangram score does not prove
The same thing no detection score proves: who wrote the text. A very low false positive rate makes a flag more informative than one from a weaker tool, and it does not change the category of evidence. The output remains a statistical judgement about how writing reads.
The only approach that records what actually happened at generation time is watermarking, and it covers a narrow slice of text — output from models whose operators chose to mark it.
And the procedural point holds regardless of which tool produced the number: a score is an input to a decision made by people. Whether an AI score should decide a misconduct case has the same answer for an accurate detector as an inaccurate one.
Where we stand, since we sell a detector
We make an AI detector. Writing a page that credits another one is uncomfortable and it is the only version of this page worth publishing, because a site that argued detection was unreliable and then quietly omitted the tool with the best evidence would be arguing, not informing.
No independent study has tested our detector. We are not in the research described above, and we are not going to imply otherwise.
Compared against
- Pangram vs GPTZero — Pangram against GPTZero on what independent testing found, who each is sold to, and what happens to the text you submit.
- Pangram vs Originality.ai — Pangram against Originality.ai on what independent testing found, who each is sold to, and what happens to the text you submit.
- Pangram vs Turnitin — Pangram against Turnitin on what independent testing found, who each is sold to, and what happens to the text you submit.
Related reading
- The accuracy research, reviewed — six studies, what each tested, and what none of them establish.
- Do AI detectors work on the newest models? — the NBER findings in short.
- How accurate AI detectors are — the wider picture, including the findings that undercut our own category.
Common questions
- Is Pangram more accurate than other AI detectors?
- On the evidence available, probably yes for medium-to-long English text — and the evidence is thinner than the headline figures suggest. An NBER working paper found it reaching "essentially zero FPRs and FNRs" on medium-to-long passages while placing GPTZero and Originality.ai in a lower tier. That paper is not peer-reviewed and says so. It is still the strongest independent signal in the category.
- What is Pangram's false positive rate?
- The company states it is "currently 1 in 10,000". That is a vendor figure about a vendor's own product, which is exactly the kind of claim this site declines to treat as established — including when the number is impressive. It is worth noting that it is stated precisely and publicly, which is more than most of the category offers.
- Does Pangram work on short text?
- Its advantage is documented on medium-to-long passages. The NBER paper's tiering was explicitly about passage length, and every detector degrades as text gets shorter because there is less signal to measure. Treat a verdict on a paragraph with much more caution than one on an essay.
- Does a Pangram result prove who wrote something?
- No, and no detector does. A very low false positive rate makes a flag more informative than it would be from a weaker tool, but the output is still a statistical judgement about how text reads, not a record of how it was produced. The only technology that records what actually happened is watermarking, which covers a narrow slice of text.
- Why does this site say detectors are unreliable and then say this one is good?
- Because both are true and the second does not rescue the first. Most tools in general use perform far worse than Pangram's published figures, and institutions run those tools rather than this one. A category being unreliable on average is compatible with one member of it being measurably better.
Sources
- 1.Artificial Writing and Automated Detection (NBER working paper, not peer-reviewed) — Jabarian & Imas — NBER Working Paper 34223, 2025
- 2.GPT detectors are biased against non-native English writers — Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press), 2023
- 3.Testing of detection tools for AI-generated text — Weber-Wulff et al. — International Journal for Educational Integrity 19:26, 2023
Vendor claims read from pangram.com on 15 August 2026.
Last reviewed 15 August 2026.