HumanFlow is not one of the two tools on this page. We sell a rewriter and run a detector of our own, so weigh this accordingly — our methodology sets out how we source what appears below and what we refuse to claim.
One independent paper tested both of these directly, which makes this the rare comparison in this category with a real head-to-head result behind it.
Last reviewed 16 August 2026 · The HumanFlow team
Why people compare these two
Both sell to organisations rather than to students, both publish precise-sounding accuracy figures, and both are pitched at buyers who will act on the output. So they land on the same shortlists — for publishers checking commissioned work, for platforms screening submissions, and increasingly for institutions.
Unusually, the comparison is not guesswork. The 2025 NBER working paper covered both, on the same corpus, at the same time. Almost no other pairing on this site can say that.
A word on why this pair gets a page when neither is a common institutional detector. Both are bought by organisations that act on the output — publishers, agencies, platforms and increasingly universities — and both publish confident accuracy figures. It is the one comparison here where a genuine head-to-head measurement exists, which makes it the best available demonstration of what independent testing adds to a vendor's own numbers.
Side by side
No prices. Neither of these is a tool most readers choose on cost, and for one of them there is no public price at all — detector pricing, where it could be captured, is on our detector pricing page.
Dimension
Pangram
Originality.ai
Who it is sold to
Enterprise and education, with a free public detector alongside.
Publishers, agencies and SEO teams first; education second and recently. Its CEO told Gizmodo in June 2024 that the company advised against academic use; by 2026 it markets an Academic Model for Educators and a Moodle plugin.
How you get at it
Free public detector. No price text is served on its site to an automated read.
No free trial. Pro is $14.95 a month, or $12.95 a month billed annually, for 2,000 credits at 100 words each.
What it needs
Its documented advantage is on medium-to-long passages; the NBER tiering was explicitly about passage length.
Not published as a word floor. The NBER paper found it struggles on short passages.
What independent testing found
Reached “essentially zero FPRs and FNRs” on medium-to-long passages in the 2025 NBER working paper, which placed GPTZero and Originality.ai in a lower tier. That paper is not peer-reviewed and says so.
Ranked second of four in the 2025 NBER working paper, in a “secondary tier” that struggles on short passages and on text run through humanizing tools. It was not among the fourteen tools in Weber-Wulff et al. (2023).
What the vendor claims
A false positive rate “currently 1 in 10,000”, stated precisely and publicly.
99%+ accuracy and false-positive rates between 0.5% and 1.5% depending on the model — published with no named corpus, no test-set size and no date.
What the vendor says about its limits
No published statement we could find.
Its terms forbid using the output as sole grounds for discipline.
What happens to your text
Not established from its published pages by us.
Its privacy policy states that unless you opt out, it may use your scan history and results to help train and improve its models. No fixed retention period is published. Institutional agreements exclude student data from general-purpose training.
Every line above is summarised from our own examination of each tool, where the studies and vendor documents behind it are quoted and linked: Pangram and Originality.ai.
What actually separates them
The NBER result is the difference. It found Pangram reaching “essentially zero FPRs and FNRs” on medium-to-long passages, and placed Originality.ai in a “secondary tier” alongside GPTZero — ranked second of four, and struggling on short passages and on text run through humanizing tools. Same paper, same test, opposite ends of it.
Length is the condition on that result and it applies to both. Pangram's documented advantage is on medium-to-long passages; the tiering was a statement about passage length before it was a statement about tools. A verdict on a paragraph from either is worth much less than one on an essay.
Their published self-assessments are worth reading side by side, because they are precise in different ways. Pangram states a false positive rate “currently 1 in 10,000” — a specific number, publicly stated, about a specific thing. Originality.ai publishes 99%+ accuracy and false-positive rates between 0.5% and 1.5% depending on the model, with no named corpus, no test-set size and no date attached. Neither is independently verified, but only one of them is checkable in principle.
Commercially they diverge. Originality.ai is a self-serve subscription — Pro at $14.95 a month, or $12.95 billed annually, for 2,000 credits at 100 words each, with no free trial — and its privacy policy states that unless you opt out it may use your scan history and results to train its models. Pangram runs a free public detector and serves no price text to an automated read, so we publish none for it rather than guess.
The buyers differ enough that the same result means different things to each, which is easy to lose in a straight accuracy comparison. A publisher checking freelance copy can absorb a false positive — the cost is an awkward conversation and a re-check. An institution acting on one imposes a misconduct process on a person. The tool with the better false-positive rate is worth more where the consequence of being wrong is heavier, and that is not where either company's marketing points.
What the evidence supports, and what it does not
The NBER working paper is the strongest independent signal in this category and it is not peer-reviewed, which the paper says itself. One working paper is a thin foundation for a claim as strong as “essentially zero”, and it should be read with the same scepticism this site applies to vendor figures. The difference is that the authors had nothing to sell.
Neither tool appears in Weber-Wulff et al. (2023), the largest peer-reviewed multi-tool evaluation, so the deepest study in the field says nothing about either. That is a real limitation on how far this comparison can be pushed.
Originality.ai's history on academic use belongs in any procurement decision. Its CEO told Gizmodo in June 2024 that the company advised against academic use and strongly recommended against disciplinary use, on false-positive grounds. By 2026 it markets an Academic Model for Educators, while its terms still forbid using the output as sole grounds for discipline.
And the ceiling applies to the winner too. A very low false-positive rate makes a flag more informative than the same flag from a weaker tool. It remains a statistical judgement about how text reads rather than a record of how it was produced.
It is worth naming what would strengthen this comparison, since it currently rests on one paper. A second independent evaluation covering both, ideally peer-reviewed and testing short passages as well as long, would settle far more than either vendor's published figures can. Until that exists, the honest summary is one working paper favouring Pangram, and two vendor claims neither of which is checkable.
Where each one came from
Originality.ai grew out of the content-marketing world and reads like it. Its natural customer runs a publication or an agency, buys copy from freelancers, and wants to know what arrived — hence bulk scanning, site scans, a Chrome extension, and credits metered in words. It also markets against a claim about Google penalising AI content, which Google has publicly denied, and that tells you what audience the marketing was written for.
Pangram came from research rather than from marketing. Published benchmarks, arXiv papers, third-party validation, and a stated false-positive rate of 1 in 10,000 given precisely enough that someone could try to falsify it. Its go-to-market is enterprise and education, where a technical buyer reads the benchmark page before the pricing page.
Those two starting points produce different relationships with evidence, which is the real subject of this comparison. One company publishes accuracy figures as marketing material with no corpus, no sample size and no date attached; the other publishes a number specific enough to be checked and submits to independent testing. Neither has been verified by us, and only one of them is structured so that it could be.
That difference shows up in what each company does when challenged. Originality.ai's response to Google publicly contradicting its penalty claim was to keep the marketing; Pangram's response to being asked for a false-positive rate was to publish one specific enough to be tested. Neither is proof of anything about the underlying models, and both are evidence about how each company will behave the next time a figure is disputed.
If one of these has produced a result about you
Neither of these is a common institutional detector, so establish first why one of them is being applied to you. If you are a freelance writer whose client ran Originality.ai over your invoice, that is a commercial dispute rather than an academic one, and the relevant fact is that a detection score is not evidence of how a document was produced.
For writers accused by a client, the strongest response is process rather than argument. Drafts, version history, research notes, the timestamps on the document. The same evidence that answers an academic accusation answers a commercial one, and it is more persuasive than any statement about detector reliability because it is about your work rather than about their tool.
If the result came from Originality.ai in an academic context, its own terms forbid using the output as sole grounds for discipline — and its CEO argued in June 2024 against academic use altogether, on the grounds that students submit too few essays for the false-positive risk to be acceptable. Both are the vendor's words, which makes them harder to dismiss than ours.
If the result came from Pangram, expect it to carry more weight and respond accordingly. The 2025 NBER working paper found it reaching essentially zero error rates on medium-to-long passages, so a flag from it is more informative than the same flag from a weaker tool. It is still a statistical judgement about how text reads rather than a record of how it was made, and it degrades on short passages like everything else in the category.
What we could not establish
The head-to-head result on this page rests on a single working paper that is not peer-reviewed and says so. One paper is thin ground for a claim as strong as essentially zero, and we would apply the same scepticism to it that we apply to vendor figures — the difference is that the authors had nothing to sell, not that the method is beyond question.
Neither tool appears in Weber-Wulff et al. (2023), the largest peer-reviewed multi-tool evaluation. The deepest study in this field says nothing about either of them, which is a real limit on how far this comparison can be pushed.
Pangram's false-positive rate of 1 in 10,000 has not been independently verified. It is stated precisely and publicly, which is more than most of the category manages, and that is a statement about its form rather than its truth.
No annual price is quoted for Pangram anywhere on this site. Its page states annual only as a saving — Save $60, Save $240 — and dividing that out to produce a rate would be arithmetic presented as a capture.
Neither company publishes the training corpus behind its detector, and that omission matters more than the missing sample sizes. What a classifier treats as ordinary prose is decided by what it was trained on, so the composition of that corpus is the best available predictor of who will be wrongly flagged. It is not published by anyone in this category, ourselves included.
Which to pick
Pick Pangram if accuracy on medium-to-long English text is the deciding factor and the cost of a wrong call is high. It is the better-performing tool in the only study that tested both, and the margin is not marginal.
Pick Originality.ai if you need the surrounding product rather than the classifier — bulk scanning, site scanning, the publisher and agency workflow it was built for — and you have read its scan-history training clause.
What neither score proves
A detector reports how statistically machine-typical a piece of prose reads. It has no access to how the text was produced, so it cannot establish authorship in either direction — a flag is not evidence of AI use, and a clean result is not a clearance. If you have been accused on the strength of one, drafting history is what actually answers it. And if your institution requires you to disclose AI assistance, disclose it — nothing on this page changes that obligation.
Pangram vs Originality.ai: common questions
Has anyone tested Pangram and Originality.ai against each other?
Yes. The 2025 NBER working paper covered both on the same corpus, finding Pangram reached essentially zero false-positive and false-negative rates on medium-to-long passages while placing Originality.ai in a secondary tier, ranked second of four. That paper is not peer-reviewed and says so, but it is the closest thing to a head-to-head this category has.
Which has the lower false positive rate?
On the independent evidence, Pangram, on medium-to-long English text. On published self-assessments, Pangram states 1 in 10,000 and Originality.ai states 0.5% to 1.5% depending on the model. Neither vendor figure is independently verified, and Originality.ai's is published without a corpus, a sample size or a date.
Does either result hold for short text?
No. The NBER paper's tiering was explicitly about passage length, and it found Originality.ai struggles on short passages specifically. Pangram's documented advantage is on medium-to-long text. Every detector degrades as passages shorten, because there is less signal to measure.
Which one costs less?
Originality.ai publishes Pro at $14.95 a month, or $12.95 a month billed annually, for 2,000 credits at 100 words each, with no free trial. Pangram serves no price text to an automated read, so we publish no figure for it rather than infer one. It does run a free public detector.
Should either be used to discipline a student?
Originality.ai's own terms forbid using its output as sole grounds for discipline, and its CEO argued in 2024 against academic use entirely on false-positive grounds. Pangram publishes no equivalent restriction. The general position holds regardless: a detector output is an input to a human decision, and drafting history carries more weight in an appeal than any percentage.