HumanFlow is not one of the two tools on this page. We sell a rewriter and run a detector of our own, so weigh this accordingly — our methodology sets out how we source what appears below and what we refuse to claim.
These two sit at opposite ends of the only independent evidence there is. On long passages one approaches zero error; the other flagged half the human writing in a peer-reviewed test set.
Last reviewed 16 August 2026 · The HumanFlow team
Why people compare these two
They are the two names that come up when someone has stopped asking which detector is cheapest and started asking which one is least likely to be wrong about them. GPTZero because it is the tool most people have heard of and the one an individual instructor is most likely to be running; Pangram because it is the tool that keeps turning up at the top of the small number of tests nobody paid for.
It is also a pair where the popular answer and the evidenced answer point in different directions, which is rarer in this category than you would hope. Most detector comparisons resolve into a coin flip between two tools nobody has tested. This one does not.
Worth setting expectations about what this page can settle. It is not a bake-off — we have not run these two against a corpus of our own, and we say so rather than implying a test we did not do. What it does is assemble what independent researchers found, what each company says about itself, and what follows for someone deciding which result to trust. That is a weaker thing than a benchmark and it is the strongest thing available without running one.
Side by side
No prices. Neither of these is a tool most readers choose on cost, and for one of them there is no public price at all — detector pricing, where it could be captured, is on our detector pricing page.
Dimension
Pangram
GPTZero
Who it is sold to
Enterprise and education, with a free public detector alongside.
Consumers. An individual instructor can sign up between classes, which means the person reading your score may have no institutional guidance on how to read it — and your institution may not know the tool is in use.
How you get at it
Free public detector. No price text is served on its site to an automated read.
Free tier of 10,000 words a month, up to 10,000 characters in a single scan.
What it needs
Its documented advantage is on medium-to-long passages; the NBER tiering was explicitly about passage length.
No published minimum. The 2025 NBER working paper found it unsuitable for very short text.
What independent testing found
Reached “essentially zero FPRs and FNRs” on medium-to-long passages in the 2025 NBER working paper, which placed GPTZero and Originality.ai in a lower tier. That paper is not peer-reviewed and says so.
Recorded the highest false-positive probability of the fourteen tools in Weber-Wulff et al. (2023), at 50% against 0% for the best in that set. The 2025 NBER working paper placed it in a “secondary tier”, unsuitable for very short text and susceptible to humanizing tools.
What the vendor claims
A false positive rate “currently 1 in 10,000”, stated precisely and publicly.
No accuracy figure is legible on its pricing page, which renders in the browser and serves $0/mo to an automated read.
What the vendor says about its limits
No published statement we could find.
No published statement we could find.
What happens to your text
Not established from its published pages by us.
Not established from its published pages by us.
Every line above is summarised from our own examination of each tool, where the studies and vendor documents behind it are quoted and linked: Pangram and GPTZero.
What actually separates them
The gap is on false positives, and it is the widest gap in the published record. In Weber-Wulff et al. (2023) — fourteen detectors, 54 test cases, 756 tests, and still the largest peer-reviewed multi-tool evaluation there is — GPTZero recorded a false-positive probability of 50%, the highest in the set, against 0% at the best. Half the human-written samples were flagged. Pangram was not in that study, but the 2025 NBER working paper that did test it reported “essentially zero FPRs and FNRs” on medium-to-long passages while explicitly placing GPTZero in a lower tier.
Length is the condition on all of it, and it cuts the same way for both. Pangram's advantage is documented on medium-to-long passages; the NBER tiering was a statement about passage length before it was a statement about tools. A verdict from either on a single paragraph is worth much less than a verdict on an essay, because there is less signal in a paragraph to read.
Who runs them differs as much as how well they work. GPTZero is a consumer product with a 10,000-word free tier that an instructor can be using within a minute of deciding to, often without institutional guidance on how to read the output and sometimes without the institution knowing. Pangram sells to enterprises and education alongside a free public detector. If you are a student, the practical question is usually not which is better but which one is pointed at you — and it is far more often GPTZero.
There is a second-order effect worth naming, because it changes who bears the risk. GPTZero's free tier means the marginal cost of running it on another student is zero, which encourages scanning everything; Pangram's is a purchase, which encourages scanning what warrants it. A tool that is free to point at people gets pointed at more people, and a false-positive rate matters in proportion to how often the tool is fired.
What the evidence supports, and what it does not
The NBER paper is the strongest independent signal in this category and it is not peer-reviewed, which the paper itself says. One working paper is a thin foundation for a claim as strong as “essentially zero”, and it deserves the same scepticism this site applies to vendor figures — the difference is that the authors had nothing to sell, not that the method is beyond question.
Pangram's own published figure is a false positive rate of “currently 1 in 10,000”. We do not treat that as established, on the same rule that keeps every other vendor's number on this site framed as a claim. It is worth noting that it is stated precisely and publicly, which most of the category does not manage.
None of this makes a Pangram flag proof. A very low false-positive rate makes a flag more informative than the same flag from a weaker tool, and it remains a statistical judgement about how text reads rather than a record of how it was produced. The only technology that records what actually happened is watermarking, and it covers a narrow slice of text.
One more asymmetry belongs in the record. GPTZero has been examined by two independent teams and Pangram by one, so the tool that looks worse has also been looked at harder. That is not a reason to discount the finding — Weber-Wulff et al. is the most thorough work in the field — but 'measured badly' and 'barely measured' are different epistemic positions, and Pangram is closer to the second than its headline figures suggest.
Where each one came from
GPTZero was released publicly in January 2023 by Edward Tian, then a Princeton undergraduate, weeks after ChatGPT arrived and months before Turnitin shipped its own indicator. It was the first detector most people ever used, and being first in a panic is a powerful distribution advantage that has very little to do with being right. The product still shows its origin: a free tier, a browser extension, a dashboard aimed at an individual teacher rather than at an institution.
Pangram arrived later and from the opposite direction — a research-credentialed team publishing benchmarks and arXiv papers, selling to enterprises and education, with the free public detector as proof rather than as the business. That ordering matters for how you read each company's claims. GPTZero grew by being the detector people had heard of; Pangram has to win procurement conversations against incumbents, which is a market where a published false-positive rate is a sales asset rather than a liability.
Neither origin makes a tool accurate. What they explain is why the evidence looks the way it does: the tool with the largest user base has the worst measured false-positive rate in the peer-reviewed literature, and the tool with the best independent showing is the one most students will never encounter.
One consequence of that timing is still visible in how each is written about. GPTZero was covered heavily by the press in 2023, when there was no independent evidence to report, so a large amount of the coverage that still ranks describes a product nobody had yet tested. Pangram arrived after the research literature existed and has been written about mostly by people citing it. Search results for these two are therefore weighted toward two different eras of what was knowable.
If one of these has produced a result about you
Establish which tool produced the number, because your position differs sharply. A GPTZero result usually means an individual instructor is running a consumer subscription — possibly with no institutional policy governing how the output is read, and possibly without the institution knowing the tool is in use at all. That is worth asking about directly and politely: which tool, what threshold, and what the department's policy says about acting on it.
If the answer is GPTZero, the Weber-Wulff finding is the most useful thing you can bring. A 50% false-positive probability across fourteen tools is not an argument that you are innocent — it is an argument that the evidence is too weak to carry a finding on its own, which is a different and much stronger claim to make in a meeting.
If the answer is Pangram, expect the number to carry more weight, and respond with process rather than with statistics. A tool with a genuinely low false-positive rate makes a flag more informative, so the productive move is producing your drafting record — version history, notes, search history, outlines. That evidence is about how the document was made, which is the question a detector cannot answer at any accuracy level.
What we could not establish
No study has tested these two against each other. Pangram's result comes from the 2025 NBER working paper and GPTZero appears in both that paper and Weber-Wulff et al. (2023), but the ranking between them here is assembled across two studies with different corpora, different years and different tool versions. That is weaker than a head-to-head and it is the best available.
We could establish nothing about what either company does with submitted text. Neither publishes a clear statement we could quote on retention or model training, which for a tool people paste unpublished work into is a real gap rather than a technicality.
Pangram's pricing was captured on 17 August 2026 and its annual rate is not quotable — the page states annual only as a saving. GPTZero's pricing could not be captured at all: its page renders in the browser and serves $0/mo to an automated read.
We have not tested either tool ourselves and this page does not pretend otherwise. Running a credible comparison would mean a pre-2022 human corpus, a matched set of model outputs, and results published in full — the original research that docs/11-programme-status.md lists as the highest-value unbuilt item on this site. Until somebody does that, every ranking in this category including ours is an assembly of other people's tests, and it should be read as one.
Which to pick
Pick Pangram if you are choosing a detector and the cost of a wrong accusation is high. On the available independent evidence it is the better-performing tool on medium-to-long English text, and the gap is not marginal.
Pick GPTZero if the choice is not yours — which is the usual case. GPTZero is what you will most often be measured by, so knowing its documented weakness on false positives is more useful than knowing a better tool exists.
What neither score proves
A detector reports how statistically machine-typical a piece of prose reads. It has no access to how the text was produced, so it cannot establish authorship in either direction — a flag is not evidence of AI use, and a clean result is not a clearance. If you have been accused on the strength of one, drafting history is what actually answers it. And if your institution requires you to disclose AI assistance, disclose it — nothing on this page changes that obligation.
Pangram vs GPTZero: common questions
Is Pangram more accurate than GPTZero?
On medium-to-long English text, the independent evidence says yes and the margin is wide. The 2025 NBER working paper reported Pangram reaching essentially zero false-positive and false-negative rates while placing GPTZero in a secondary tier. Separately, the largest peer-reviewed multi-tool test put GPTZero's false-positive probability at 50%, the worst of fourteen. Two caveats matter: the NBER paper is not peer-reviewed, and neither result transfers to short passages.
Why is GPTZero so widely used if it tested badly?
Because distribution and accuracy are unrelated. GPTZero launched in January 2023, weeks after ChatGPT, and became the default consumer detector before any independent testing existed. It is free at 10,000 words a month, requires no institutional purchase, and an individual instructor can adopt it alone. Pangram sells to organisations. Nothing about being the tool people reach for implies being the tool that is right.
If Pangram flags my essay, does that prove I used AI?
No. A lower false-positive rate makes the flag more informative, not conclusive. Every detector outputs a statistical judgement about how text reads, and predictable, conventional prose reads as machine-like whoever wrote it. What a flag should trigger is a look at your drafting record, not a finding.
Does either tool work on short text?
Both degrade, and the evidence for Pangram's advantage does not extend there. The NBER paper's tiering was explicitly about passage length, and it found GPTZero unsuitable for very short text specifically. Treat any verdict on a paragraph from either tool with far more caution than one on a full essay.
Which one would a university be running?
Most likely neither at an institutional level — the detectors bought by universities are Turnitin and Copyleaks. GPTZero turns up as an individual instructor's own subscription, which is a different situation with different implications, because there may be no institutional policy governing how the result is read.