humanflow

Pangram vs Turnitin

HumanFlow is not one of the two tools on this page. We sell a rewriter and run a detector of our own, so weigh this accordingly — our methodology sets out how we source what appears below and what we refuse to claim.

Turnitin is best of fourteen in the strongest study; Pangram is best of four in the newest one. The two results cannot be combined, and the studies are not equally strong.

Last reviewed 16 August 2026 · The HumanFlow team

Why people compare these two

This is the comparison an institution reaches when it already has Turnitin and someone asks whether something better exists. It is a live procurement question rather than a hypothetical, because Pangram is one of the few detectors positioned at organisations rather than individuals.

It is also a clean test of how to weigh evidence in this field, because the two tools won different studies of different sizes and vintages, and the temptation to read that as a ranking is strong and wrong.

This is a procurement page more than a consumer one, and it is worth saying what that means for how it is written. Nobody is choosing between these two on a Tuesday afternoon; a committee is deciding whether to disturb something that already works. So the comparison below weighs switching cost and evidential strength against each other rather than presenting a winner, because that is the actual shape of the decision.

Side by side

No prices. Neither of these is a tool most readers choose on cost, and for one of them there is no public price at all — detector pricing, where it could be captured, is on our detector pricing page.

DimensionPangramTurnitin
Who it is sold toEnterprise and education, with a free public detector alongside.Institutions only. A student or an individual instructor cannot buy it; it arrives inside the systems a university already licenses.
How you get at itFree public detector. No price text is served on its site to an automated read.No public checker. If your institution has not enabled the indicator for your class, there is no way to see your own score.
What it needsIts documented advantage is on medium-to-long passages; the NBER tiering was explicitly about passage length.At least 300 words of prose, up to 30,000. .docx, .pdf, .txt or .rtf under 100MB, in English, Spanish or Japanese.
What independent testing foundReached “essentially zero FPRs and FNRs” on medium-to-long passages in the 2025 NBER working paper, which placed GPTZero and Originality.ai in a lower tier. That paper is not peer-reviewed and says so.Scored highest of the fourteen tools in Weber-Wulff et al. (2023) — in a study whose own conclusion was that detection tools “are neither accurate nor reliable”, with every tool below 80% accuracy.
What the vendor claimsA false positive rate “currently 1 in 10,000”, stated precisely and publicly.Aims to keep false positives under 1% above a 20% detected share. Below 20% it attributes no score at all, showing an asterisk, because its own testing found more false positives in that band.
What the vendor says about its limitsNo published statement we could find.States its indicator is not intended as the sole basis for an academic misconduct finding.
What happens to your textNot established from its published pages by us.Submissions may be retained in its repository depending on the institution's configuration, which the student does not set.

Every line above is summarised from our own examination of each tool, where the studies and vendor documents behind it are quoted and linked: Pangram and Turnitin.

What actually separates them

Incumbency is the real asymmetry. Turnitin is already inside the systems a university licenses, and the AI indicator arrived in a product institutions had bought for similarity checking years earlier. Replacing it is not a like-for-like swap; it is a change to submission workflow, instructor training and appeal procedure. Pangram sells to enterprise and education alongside a free public detector, and serves no price text to an automated read, so we publish no figure for it.

Turnitin's product decisions around uncertainty have no equivalent on Pangram's side. It attributes no score at all between 0% and 20%, showing an asterisk, because its own testing found more false positives in that band, and it aims to keep false positives under 1% above that threshold. It also states that its indicator is not intended as the sole basis for a misconduct finding. Pangram publishes a false positive rate of “currently 1 in 10,000” and no equivalent restriction we could find.

Coverage differs in ways that decide some deployments outright. Turnitin needs 300 to 30,000 words of prose in .docx, .pdf, .txt or .rtf under 100MB, in English, Spanish or Japanese. Pangram's documented strength is on medium-to-long passages, and the tiering that produced it was explicitly a statement about passage length — so short-answer assessment is a weak setting for either, and a stated three-language limit is a hard constraint for a multilingual institution.

For a student the difference is smaller than it looks. Neither is something you can run on your own work: Turnitin has no public checker and, unless your institution enables the indicator for your class, no route to your own score. Pangram's free public detector is a different matter, but a result from it is not the result your institution will see.

Switching costs are the part of this comparison a benchmark cannot price, and they are not merely inconvenience. An institution's appeals policy names the tool, its instructors are trained on one report format, and its historical cases were decided on one scale. Changing detector means every one of those has to move too, and a department part-way through that transition is running two standards at once — which is worse for students than either standard alone.

What the evidence supports, and what it does not

Turnitin scored highest of the fourteen tools in Weber-Wulff et al. (2023) — 54 test cases, 756 tests, peer-reviewed, and the largest multi-tool evaluation in the field. Pangram was not in it. Its own result comes from the 2025 NBER working paper, which found it reaching “essentially zero FPRs and FNRs” on medium-to-long passages while placing GPTZero and Originality.ai a tier below. Turnitin was not in that one.

The two cannot be combined into a ranking, and it is worth being explicit about why rather than just asserting it. Different corpora, different years, different tool versions, different scoring methods — and one is peer-reviewed while the other is a working paper that says of itself that it is not. A newer result on a smaller field is not automatically a stronger one.

What both studies agree on is the ceiling. Weber-Wulff and colleagues concluded that detection tools “are neither accurate nor reliable”, with every tool below 80% accuracy, and measured 26% accuracy across all fourteen on machine-paraphrased text. Pangram's near-zero figures are explicitly conditioned on medium-to-long passages. Neither result licenses treating an output as proof, and Pangram's own strong showing does not change what a flag is: a statistical judgement about how text reads, not a record of how it was produced.

Both figures rest on studies that excluded the other tool, so the comparison is structurally incomplete in a way no amount of careful reading fixes. What can be said is narrower and still worth having: Turnitin is the most thoroughly evaluated detector in the peer-reviewed literature, and Pangram is the strongest performer in the newest independent work. Those are compatible statements about different evidence, not competing claims about one thing.

Where each one came from

Turnitin's position in this market was won before AI detection existed. Two decades of similarity checking built the integrations, the procurement relationships and the institutional habits, and the AI indicator inherited all of it — which is why it is the most widely deployed detector in education without ever having won a comparison on detection.

Pangram has no such inheritance and has competed on the only ground available to a newcomer: published evidence. Benchmarks, arXiv papers, third-party validation, a stated false-positive rate of 1 in 10,000. That is the strategy of a company that has to displace an incumbent rather than defend a base, and it produces exactly the kind of material a technical buyer can evaluate.

So the comparison is incumbency against evidence, which is a genuine institutional dilemma rather than a rhetorical one. Turnitin's AI indicator is already installed, already integrated, already in the appeals policy. Replacing it is not a like-for-like swap but a change to submission workflow, instructor training and disciplinary procedure — costs a benchmark result does not address and a procurement committee cannot ignore.

If one of these has produced a result about you

In practice this will almost always be Turnitin, because that is what institutions run. Start by asking which band the score falls in: Turnitin attributes no score at all between 0% and 20% and displays an asterisk instead, because its own testing found more false positives in that range. An asterisk is a refusal to give a number and is regularly mistaken for a low one.

Then check the length and the language of what was submitted. Turnitin requires at least 300 words of prose and covers English, Spanish and Japanese; a result on a shorter document, or on writing in another language, rests on conditions the tool's own documentation does not support.

Turnitin states that its indicator is not intended as the sole basis for an academic misconduct finding. That sentence is the most useful thing you can bring to a meeting, because it comes from the company selling the tool and it describes how the output is meant to be used rather than how reliable it is — a claim nobody in the room can dispute on your behalf.

If Pangram produced the result, the response is the same in structure and different in emphasis. Its independent showing is genuinely strong, so arguing about detector reliability in general will not get far. Produce the drafting record instead, and note that Pangram's documented advantage is on medium-to-long passages — a verdict on a short answer is outside the conditions its published performance was measured under.

What we could not establish

No study has tested these two against each other, and the two results usually quoted are not combinable. Turnitin was best of fourteen tools in Weber-Wulff et al. (2023), which is peer-reviewed and did not include Pangram. Pangram reached essentially zero error rates in the 2025 NBER working paper, which is not peer-reviewed and did not include Turnitin.

It is worth being explicit about why that cannot be resolved by preferring the newer result. Different corpora, different years, different tool versions, different scoring methods, and one peer-reviewed while the other is a working paper that says of itself that it is not. A newer result on a smaller field is not automatically a stronger one.

Pangram's 1 in 10,000 false-positive rate is unverified by anyone but Pangram, and Turnitin's aim of under 1% is unverified by anyone but Turnitin. Neither company has published the corpus behind its number.

Turnitin's pricing is not public and its documentation returns 403 to us. Pangram's prices were read on 17 August 2026, and no annual rate is quoted for it because its page states annual only as a saving rather than as a figure.

We could not establish whether any institution has actually replaced Turnitin's indicator with Pangram, or what happened if one did. A published migration — what changed in flag rates, in appeals, in instructor confidence — would be worth more to a procurement committee than either vendor's benchmark, and as far as we can find nobody has written one.

Which to pick

Pick Pangram if you are evaluating detectors on evidence rather than on installed base, your assessments are medium-to-long English prose, and you can absorb the workflow change. It is the better performer in the only study that measured it.

Pick Turnitin if you already have it, you assess in Spanish or Japanese as well as English, or you value a vendor that withholds a score in the band where its own testing says the score is least trustworthy.

What neither score proves

A detector reports how statistically machine-typical a piece of prose reads. It has no access to how the text was produced, so it cannot establish authorship in either direction — a flag is not evidence of AI use, and a clean result is not a clearance. If you have been accused on the strength of one, drafting history is what actually answers it. And if your institution requires you to disclose AI assistance, disclose it — nothing on this page changes that obligation.

Pangram vs Turnitin: common questions

Is Pangram more accurate than Turnitin?
No study has tested both, so there is no measurement that answers this. Turnitin was best of fourteen tools in Weber-Wulff et al. (2023), which is peer-reviewed and did not include Pangram. Pangram reached essentially zero error rates on medium-to-long passages in the 2025 NBER working paper, which is not peer-reviewed and did not include Turnitin.
Should an institution switch from Turnitin to Pangram?
Not on the accuracy evidence alone, because that evidence does not compare them. The considerations that can be settled are practical: language coverage, where Turnitin's indicator handles English, Spanish and Japanese; passage length, where both weaken on short text; and the workflow, training and appeals changes any swap entails.
Can a student run either one on their own work?
Turnitin, no — it sells to institutions only, publishes no self-serve checker, and unless your institution has enabled the indicator for your class there is no route to your score. Pangram runs a free public detector, but a result from it is not what your institution will see and should not be treated as a preview.
What is Pangram's published false positive rate?
The company states it is currently 1 in 10,000. That is a vendor figure about a vendor's own product, which this site declines to treat as established — including when the number is impressive. It is worth noting that it is stated precisely and publicly, which most of this category does not manage.
Does either work on short answers?
Both weaken, and Pangram's advantage in particular is documented on medium-to-long passages rather than short ones. Turnitin sets a hard floor of 300 words for long-form writing, having raised it from 150 because accuracy improves with more text. Short-answer assessment is a poor setting for detection generally.

Related