humanflow

GPTZero vs Copyleaks

HumanFlow is not one of the two tools on this page. We sell a rewriter and run a detector of our own, so weigh this accordingly — our methodology sets out how we source what appears below and what we refuse to claim.

The difference that matters most is not accuracy. It is that one of these is bought by your institution and the other by one instructor acting alone.

Last reviewed 16 August 2026 · The HumanFlow team

Why people compare these two

These are the two detectors a student is most likely to actually be measured by outside Turnitin — Copyleaks because universities buy it, GPTZero because individual instructors do. They are rarely a choice anybody makes side by side; they are two situations you can find yourself in, and they have different consequences.

They are also the pair where the published evidence points hardest in one direction while usage points in the other, which makes the comparison worth having rather than academic.

The question people actually arrive with is usually narrower than which is better — it is whether the result in front of them should count. That depends less on the classifiers than on how each tool got into the room, which is why this page spends most of its length on procurement route, policy and recourse rather than on accuracy figures that cannot be compared anyway.

Side by side

No prices. Neither of these is a tool most readers choose on cost, and for one of them there is no public price at all — detector pricing, where it could be captured, is on our detector pricing page.

DimensionGPTZeroCopyleaks
Who it is sold toConsumers. An individual instructor can sign up between classes, which means the person reading your score may have no institutional guidance on how to read it — and your institution may not know the tool is in use.Institutions and individuals both — one of the few detectors genuinely sold to universities rather than only to people. Education is priced on full-time student count, with integrations for Canvas, Moodle, D2L Brightspace, Blackboard, Schoology, Sakai and Edsby.
How you get at itFree tier of 10,000 words a month, up to 10,000 characters in a single scan.Paid. Personal is $16.99 a month, or $13.99 a month billed annually; one credit covers up to 250 words.
What it needsNo published minimum. The 2025 NBER working paper found it unsuitable for very short text.Not published as a word floor. Its highest sensitivity setting is described as designed to flag text put through a humanizer or spinner.
What independent testing foundRecorded the highest false-positive probability of the fourteen tools in Weber-Wulff et al. (2023), at 50% against 0% for the best in that set. The 2025 NBER working paper placed it in a “secondary tier”, unsuitable for very short text and susceptible to humanizing tools.Best of the seven tools in Perkins et al. (2024) — and it still missed 39% of the AI cases and produced the highest false-accusation rate in that set. Its accuracy fell from 73.9% to 58.7% once simple adversarial techniques were applied.
What the vendor claimsNo accuracy figure is legible on its pricing page, which renders in the browser and serves $0/mo to an automated read.99.97% accuracy and a 0.026% false-positive rate on its V10 model.
What the vendor says about its limitsNo published statement we could find.No published statement we could find.
What happens to your textNot established from its published pages by us.Paid plans include the Shared Data Hub, which Copyleaks describes as a library of user-submitted documents that scans are compared against. Its privacy policy states it uses this information to train its models, with opt-out available to direct customers by contacting support.

Every line above is summarised from our own examination of each tool, where the studies and vendor documents behind it are quoted and linked: GPTZero and Copyleaks.

What actually separates them

Start with who bought it, because it decides what recourse you have. Copyleaks is sold to institutions on full-time student count, with integrations for Canvas, Moodle, D2L Brightspace, Blackboard, Schoology, Sakai and Edsby. If it flagged you, there is a procurement decision, an institutional policy and usually an appeals process behind that flag. GPTZero costs an instructor nothing to start — 10,000 words a month free, up to 10,000 characters a scan — and can be in use in a classroom without the institution knowing, which means there may be no policy governing how the number is read.

The published record on false positives separates them sharply. In Weber-Wulff et al. (2023), GPTZero recorded the highest false-positive probability of the fourteen tools examined, at 50% against 0% for the best in that set. Copyleaks was not in that study; in Perkins et al. (2024) it was the best of seven — while still producing the highest false-accusation rate of those seven, which is a reminder that “best available” and “safe to act on” are different claims.

On adversarial text they diverge in what they claim rather than in what has been shown. Copyleaks's highest sensitivity setting is explicitly described as designed to flag text run through a humanizer or spinner. GPTZero was found susceptible to humanizing tools by the 2025 NBER working paper. Neither should be relied on there: across all fourteen tools, Weber-Wulff and colleagues measured 26% accuracy on machine-paraphrased text.

What happens to your text differs most of all, and only one side of it is documented. Copyleaks's paid plans include the Shared Data Hub, a library of user-submitted documents that scans are compared against, and its privacy policy states it uses this information to train its models, with opt-out available to direct customers by contacting support. For GPTZero we could establish nothing either way from its published pages, which is a different thing from a denial.

Consider what each arrangement means when the tool is wrong, because that is the scenario worth planning for. A Copyleaks false positive lands inside a system with a policy, an appeals route and a named owner. A GPTZero false positive from an individual subscription may land nowhere at all — there may be no policy covering the tool, no record that it was used, and no route of appeal that acknowledges the result exists.

What the evidence supports, and what it does not

The two have never been tested against each other, so any ranking between them is assembled from studies that share no corpus. What can be said is narrower and more useful: GPTZero has a measured false-positive problem in the largest peer-reviewed test of these systems, and Copyleaks does not appear in that test at all.

Copyleaks's own figures — 99.97% accuracy and a 0.026% false-positive rate on its V10 model — sit about forty points above what independent testing found under adversarial conditions, where its accuracy fell from 73.9% to 58.7%. That gap is not evidence of dishonesty. It is what happens when a vendor measures undisguised model output and a researcher measures text somebody tried to disguise.

GPTZero publishes no accuracy figure we could read. Its pricing page renders in the browser and serves $0/mo to an automated request, which is why it appears in no captured price table on this site.

The two tools are also tested against different threat models, which is easy to miss when comparing headline results. Perkins et al. subjected Copyleaks to deliberate adversarial techniques and measured the drop, from 73.9% to 58.7%. Weber-Wulff et al. measured GPTZero largely on the false-positive side, on human writing that was not trying to evade anything. Those are different questions, and neither study answers the other's.

Where each one came from

Copyleaks sells content integrity to organisations — universities priced on full-time student count, and enterprises with compliance requirements. Everything about how it is bought follows from that: LMS integrations for Canvas, Moodle, D2L Brightspace, Blackboard, Schoology, Sakai and Edsby, a procurement conversation, and a contract. Nobody adopts Copyleaks quietly.

GPTZero was built to be adopted quietly, and that is not a criticism of its engineering. A free tier of 10,000 words a month, no purchase order, no IT involvement — an instructor who read a worrying article at breakfast can be running it by the first class. In January 2023 that accessibility was the entire reason the tool mattered, because the institutional options did not exist yet.

Three years later the same property is the problem this page is really about. A detector inside a procurement process arrives with a policy attached; a detector adopted by one person arrives with whatever that person believes about it. The difference in how a flagged student is treated has almost nothing to do with the classifiers and almost everything to do with which of those two routes the tool took into the room.

If one of these has produced a result about you

Find out which tool it was, because your available moves are genuinely different. That question is reasonable to ask directly and is not an accusation: knowing whether a result came from an institutionally licensed system or an individual subscription tells you which policy applies.

If it was Copyleaks, there is an institutional process behind it. Ask for the policy in writing, ask which sensitivity setting was used — the highest is explicitly designed to flag humanized and spun text and accepts more false positives to do it — and ask what weight the department gives the number. A procurement decision means somebody has already had to justify using it.

If it was GPTZero, ask whether your institution has a policy on AI detection at all, and whether this tool is covered by it. It is a common situation and not a hostile question. Where no policy exists, the department is making an evidential judgement without institutional guidance, and the peer-reviewed record on this specific tool is directly relevant: the highest false-positive probability of the fourteen tools in Weber-Wulff et al. (2023), at 50%.

Whichever it was, the answer is a drafting record rather than a counter-argument about statistics. Version history, notes, outlines, and the ability to talk about your own argument in a room. Detectors measure how predictable prose is; none of them observes how a document was written, which is the actual question.

What we could not establish

These two have never been tested against each other. GPTZero's false-positive figure comes from Weber-Wulff et al. (2023), which did not include Copyleaks; Copyleaks's ranking comes from Perkins et al. (2024), which did not include GPTZero. Any ordering between them is assembled across studies and is weaker than it will look in a table.

We could establish nothing about GPTZero's handling of submitted text from its published pages — no retention period, no statement on training use, in either direction. Copyleaks is the opposite case: the Shared Data Hub and the training clause are documented, and that documentation is what makes them criticisable.

GPTZero's pricing could not be captured. Its page renders in the browser and serves $0/mo to an automated read, which is the worked example that docs/10-competitor-capture.md exists for.

Copyleaks's figures were read on 15 August 2026 and could not be re-verified two days later — the page now returns a Cloudflare challenge. They stand as captured and need a browser read to confirm.

Which to pick

Pick GPTZero if you need a free check on your own writing before submitting and you will treat the answer as a rough signal. Know its documented weakness on false positives before you let a flag change work you wrote yourself.

Pick Copyleaks if you are the institution. It is the tool with a real integration story, a documented position on humanized text and independent testing behind it — provided you have read what the Shared Data Hub does with student submissions before you turn it on.

What neither score proves

A detector reports how statistically machine-typical a piece of prose reads. It has no access to how the text was produced, so it cannot establish authorship in either direction — a flag is not evidence of AI use, and a clean result is not a clearance. If you have been accused on the strength of one, drafting history is what actually answers it. And if your institution requires you to disclose AI assistance, disclose it — nothing on this page changes that obligation.

GPTZero vs Copyleaks: common questions

Which is more accurate, GPTZero or Copyleaks?
No study has tested both. What the record does show is that GPTZero recorded the highest false-positive probability of fourteen tools in Weber-Wulff et al. (2023), at 50%, and that Copyleaks was the best of seven in Perkins et al. (2024) while still producing the highest false-accusation rate in that smaller set. Those are different studies with different corpora, so they do not combine into a ranking.
Can an instructor use GPTZero without the university knowing?
Yes, and this is the practical difference between the two. GPTZero is a consumer product with a free tier that an individual can start using immediately. Copyleaks is bought institutionally and integrated into the LMS. If you have been flagged by a tool your institution did not procure, whether any policy governs how that result is used is a fair question to ask.
Does either one catch text that has been through a humanizer?
Copyleaks claims to, via its highest sensitivity setting. GPTZero was found susceptible to humanizing tools by the 2025 NBER working paper. Neither is reliable there — Weber-Wulff and colleagues measured 26% accuracy across fourteen detectors on machine-paraphrased text, which is the finding that recurs across the whole literature.
Is my work stored if my university runs Copyleaks?
Copyleaks's paid plans include the Shared Data Hub, which it describes as a library of user-submitted documents that scans are compared against, and its privacy policy states it uses this information to train its models. Opt-out is available to direct customers by contacting support. What your institution has configured is a question for your institution.
Should I check my essay in GPTZero before handing it in?
You can, and it will tell you less than you want. It is a different classifier from whatever your institution runs, and detectors disagree with each other on identical text — so neither a clean result nor a flag tells you what will happen. The more useful preparation is a drafting record: version history, notes and outlines are what actually answer an accusation.

Related