humanflow

Sapling's AI detector

Sapling publishes a false positive rate — under 3% — which almost nobody in this category does. It is the honest number to publish and it is not a small one. Across a 300-student cohort that is roughly nine people wrongly flagged.

Last reviewed 16 August 2026 · The HumanFlow team

The number they publish, and the one they do not have to

Read from their detector page on 16 August 2026: a 97%+ detection rate for AI-generated content, and less than 3% false positives for human-written content, both measured on their own benchmarks using longer texts.

The second figure is the one worth crediting. A false positive rate measures how often the tool accuses a person; a detection rate measures how much machine text it catches. The first is flattering and almost everyone publishes it. The second is the number that decides whether a student gets an email, and Sapling is one of very few vendors that states it at all.

What under 3% looks like from the other end

Percentages this small stop feeling small the moment you multiply them by a class list. On a 300-student cohort submitting one piece each, a 3% rate is about nine people wrongly flagged. Over a three-module term it is closer to twenty-seven flags, spread across students who did nothing.

None of that makes the figure bad. It is better than several published rates in this category and vastly better than the 61% that seven detectors averaged on TOEFL essays in the peer-reviewed testing. The point is that a reassuring-sounding rate and a large number of wrongly accused people are the same fact described at two scales, and institutions make decisions at the second scale while reading the first.

It also carries the qualifier their own page attaches: measured on longer texts, with shorter texts and certain content types behaving differently. A rate measured on essays does not describe what happens to a discussion post.

What the vendor says about its own tool

This is quoted rather than summarised because the wording matters. Their documentation states that no current AI content detector, including Sapling's, should be used as a standalone check to determine whether text is AI-generated or written by a human, and that false positives and false negatives will occur.

That is a stronger statement than most critics of this category make, published by a company that sells one. Their FAQ also carries the question “I pasted my own writing but it says it's AI-generated” — a vendor documenting the experience of being wrongly flagged by its own product.

If an institution is running Sapling and treating its output as sufficient, the vendor disagrees with the institution. That is a usable sentence in an appeal, and it is on the vendor's own site rather than ours.

The free tier and the length problem

Free queries are capped at 2,000 characters, roughly 350 words. The published accuracy figures come from benchmarks on unspecified “longer texts”, and their page separately notes the detector becomes much more accurate after about 50 words.

So the free tool operates in a band the headline numbers were not measured on. That is not a trick — the caveat is stated on the same page — but anyone quoting the 3% figure about a result from a short free query is quoting a number that was measured somewhere else. The general version of this is on whether text length affects detection.

What a Sapling score does not prove

That anybody used AI. Their mechanism is a Transformer assigning each token a probability of being machine-generated, with per-sentence perplexity available in the interface. It is measuring predictability, which is a property of prose rather than a fact about its author, and predictable writing has many entirely human causes.

If a score has been put to you, the evidence that helps is a drafting record, and the appeal page covers what to ask for and in what order — starting with the vendor's own statement that this is not a standalone check.

Where we stand

We sell a detector too, and we publish no false positive rate for it. Sapling publishes one and we do not, which is a gap on our side rather than theirs, and the reason is on our methodology page: we have not run a test we would be willing to stand behind. That is an explanation rather than an excuse.

Compared against

  • Sapling vs GPTZero Sapling against GPTZero on what independent testing found, who each is sold to, and what happens to the text you submit.

Figures read from Sapling's AI detector page on 16 August 2026: sapling.ai/ai-content-detector. Quoted as their claims, not independently verified here.

Common questions

What is Sapling's false positive rate?
Their page states less than 3% for human-written content, measured on their own benchmarks using longer texts, alongside a 97%+ detection rate for AI-generated content. They also note that shorter texts or certain content types may have different accuracy rates.
Is under 3% good?
It is better than most of this category publishes, and it is not small. Across a 300-student cohort submitting one essay each, a 3% rate is around nine people wrongly flagged. Whether that is acceptable depends entirely on what happens to the nine, which is a policy question rather than a technical one.
Does Sapling say its detector can be used on its own?
It says the opposite, in unusually plain terms: no current AI content detector, including Sapling's, should be used as a standalone check, and false positives and false negatives will occur. That is the vendor's own sentence about its own product and the whole category.
How much text does Sapling need?
Their page says the detector becomes much more accurate after 50 or so words, and that its published figures come from benchmarks using longer texts. Free queries are capped at 2,000 characters, roughly 350 words, so the free tool works on shorter text than the headline figures were measured on.
What does Sapling actually measure?
Their description is a Transformer that outputs, for each token, the probability that it was machine-generated — and the interface can show per-sentence perplexity. That is the standard mechanism in this category rather than anything unusual, and it is the reason predictable human writing scores badly.
Do you compete with Sapling?
Barely. They sell developer tooling, grammar and autocomplete alongside the detector; we sell a rewriter and a detector. The overlap is small enough that this page carries less conflict than most on this site, and every figure is theirs and linked regardless.