What the company claims
Read from their AI detector page on 16 August 2026. A 99% detection rate, attributed to independent evaluations from the RAID benchmark. Coverage of content from ChatGPT, Gemini, Claude and Copilot. Multilingual detection. A score from 0 to 100 expressing the likelihood that text was AI-generated, with an explainer card giving the reasons behind a flag.
A detection rate is not an accuracy figure and it is certainly not a false positive rate. It answers how much machine text the tool catches. It says nothing about how often it flags a person, which is the only number that matters to somebody defending their own work — and QuillBot does not publish it. Nor does anyone else in this category, ours included.
The disclosure worth reading twice
Their documentation states that when the result is unclear, the model tends to classify texts as human-written, which reduces false positives. That is a threshold decision stated openly, and almost nothing else in this market states one.
It has a direct consequence for how the output should be read. If ambiguous cases are resolved towards human, then a human verdict from this tool is weak evidence — it includes everything the model was unsure about. An AI verdict, by the same logic, is comparatively stronger. Most readers assume the two carry equal weight, and on this tool they do not.
It is also the honest trade. Every detection threshold buys fewer false accusations at the cost of more missed machine text, and there is no setting that avoids both. QuillBot has chosen the direction that protects the person being scored, and said so.
What they say their tool gets wrong
Their page names two categories: heavily paraphrased writing, and highly formulaic content such as legal jargon and academic writing. It also notes that longer texts produce more reliable scores than short ones.
Sit with the second category for a moment. A detector vendor is stating that conventional academic prose is more likely to be misread by its own tool — which is precisely the finding of every independent study of this problem, arriving here from the company selling the detector rather than from a critic.
What a QuillBot score does not prove
That anyone used AI. QuillBot's own documentation says the tool is not perfect and that you should never rely on an AI detector alone when reviewing results. We would put it slightly more strongly — a score is a statistical judgement about how text reads, not a record of how it was produced — but the direction is the same and it is their sentence, not ours.
If a QuillBot result has been quoted at you, the useful ground is the evidence that actually helps, which is drafting history rather than a counter-score, and what a score can support procedurally.
Where we stand, since we sell against them
QuillBot competes with our humanizer, not with our detector, and we have written about their paraphrasing suite elsewhere. That is a conflict, and the mitigation is that every figure above is theirs and linked.
On the substance we have no complaint to register. A vendor publishing a minimum text length, naming the writing its tool struggles with, disclosing which way it resolves ambiguity, and telling users not to rely on it alone is behaving better than most of this category. The gap we would still press on is the false positive rate, which they do not publish — and neither do we.
Compared against
- Scribbr vs QuillBot — QuillBot against Scribbr on what independent testing found, who each is sold to, and what happens to the text you submit.
- Copyleaks vs QuillBot — QuillBot against Copyleaks on what independent testing found, who each is sold to, and what happens to the text you submit.