humanflow
AI detection · The HumanFlow team · 12 min read

Scribbr AI Detector Review: The Academic Tool That Publishes Its Own Report Card

Scribbr's free AI detector, reviewed: 1,200-word checks, no sign-up, and the rare vendor that publishes research ranking its own tool at 84%.

Scribbr's AI detector is a free, no-sign-up checker — 1,200 words per scan, unlimited scans — from an academic-services company that does something almost no detector vendor does: it publishes comparative research grading its own tool, and the grade isn't 99%. Scribbr's own study puts its premium detector at 84% and its free one at 78%.

If you've read many detector reviews, you know how strange that paragraph is. The category's marketing default is a two-digit number starting with 9 and no methodology in sight. Scribbr instead maintains a public test in which its free tool ties with a competitor's, its premium tool catches only 60% of the hardest text category, and the write-up states outright that "AI detectors shouldn't be treated as absolute proof." This review takes that self-assessment seriously — crediting it, citing it, and pushing on its weak points.

Disclosure first: we build a detector and humanizer ourselves at HumanFlow, which makes this a competitor's review. Here's our editorial policy — judge accordingly. Everything below comes from Scribbr's pages fetched during writing, from its published research, or carries a verification flag.

Where the detector sits in Scribbr's world

Scribbr is not a detection company. It's an academic-services company — proofreading, editing, citation generation, plagiarism checking — that has spent a decade serving students and researchers, and its knowledge base of citation and methodology guides ranks all over academic search queries. The AI detector joined a suite whose center of gravity is helping students finish papers, not policing them. Scribbr's plagiarism checker runs on Turnitin's engine, per Scribbr's own product pages, which tells you how the company operates: partner or build where needed, wrap it in student-facing service.

That positioning shapes the detector's design choices. It's aimed at a student checking their own work before submission — the person who used ChatGPT to outline, wrote the draft themselves, and now wants to know whether a detector will misread them — rather than at an integrity office hunting for cheaters or a content agency scanning freelancers at scale. No API, no bulk dashboard, no institutional reporting. A paste box and a verdict.

The free tool: scope and specifics

Fetched from scribbr.com/ai-detector at review time: up to 1,200 words per submission, unlimited checks, completely free, no account required. Scribbr also states that submissions "remain private; we do not store or share your data" — a meaningful claim for students pasting unpublished thesis chapters, though as with any vendor privacy claim, it's a promise you're trusting rather than verifying.

Results come back at paragraph level, sorting text into four categories: AI-generated, AI-generated and AI-refined, human-written and AI-refined, and human-written. That mixed-content vocabulary matters. Most real documents in 2026 aren't purely one thing — a human draft tightened by an assistant, an AI outline fleshed out by a person — and a tool that only answers "AI or not" misreads the majority case. Scribbr's free version targets older model families (GPT-2, GPT-3, GPT-3.5) with what the company itself calls "average accuracy"; the premium version claims "high accuracy" and adds GPT-4-class detection. Language support covers English, Spanish, German, and French.

Read those model names again, though. The free tool's stated detection targets top out at GPT-3.5-era output, with GPT-4 reserved for premium — and the page names GPT-2 and GPT-3-class models for the free tier and adds GPT-4 for premium, claiming coverage of no model newer than that for either. Competitors' pages name current models explicitly; Scribbr's specificity about older models is more honest than a vague "detects all AI," but it leaves a real question about the newest generators. We flag it rather than guess.

Scribbr FreeScribbr Premium
Words per scan1,200Higher cap, unstated
ScansUnlimitedUnlimited
Sign-up requiredNoYes
Stated model coverageGPT-2/3/3.5, "average accuracy"Adds GPT-4, "high accuracy"
Result granularityParagraph-level, 4 categoriesSame, higher claimed accuracy
Self-published test score78%84%
Price$0Not published on the detector pages

Premium pricing wasn't published on the pages we fetched, so we won't quote a number from memory — if a figure matters to you, check the live page.

The report card: Scribbr's published research, credited properly

Here is the thing that makes Scribbr worth writing about. The company maintains a comparative study — "Best AI Detectors," last revised July 21, 2026, per the page — that tests twelve detectors, publishes the methodology, and reports results in which Scribbr does not sweep the field.

The method, briefly: 30 texts across six categories of five each — human-written, GPT-3.5-generated, GPT-4-generated, human-plus-GPT-3.5 hybrids, GPT-3.5 output paraphrased by QuillBot, and human writing paraphrased by QuillBot — scored on how close each detector's percentage came to ground truth (full credit within 15 points, half credit within 40). The results, from Scribbr's own table: Scribbr premium 84%, QuillBot and Scribbr free tied at 78%, Originality.ai 76%, Sapling 68%, Copyleaks 66%, ZeroGPT 64%, GPT-2 Output Detector and CrossPlag 58%, GPTZero 52%, Writer 38%, OpenAI's retired classifier 38%.

Three things deserve explicit credit. Scribbr ranks a direct competitor's free tool dead even with its own — QuillBot, whose detector we review separately — and ranks several household names far below both. It reports the ugly subgroup finding: even its winning premium tool caught only 60% of paraphrased or mixed texts, the category closest to how contested real-world documents actually look. And it states the conclusion its own sales page has to live with: "AI detectors can never provide 100% accuracy" and "shouldn't be treated as absolute proof." A vendor publishing numbers like 84%, 78%, and 60% about itself, while the rest of the market prints 99%, is doing the category a service. We say that as a competitor with no incentive to say it.

Now the pushback, because credit without scrutiny is just marketing by other means. Thirty texts is a small sample — five per category means one odd result swings a tool's score by huge margins, and the difference between 78% and 76% is statistically meaningless at that size. The generators tested trail the current frontier even in the July 2026 revision. The scoring rubric — distance-based partial credit — is defensible but nonstandard, and it isn't directly comparable to the sensitivity/specificity framing used in academic evaluations like the RAID benchmark. And there's an obvious structural concern when the test's author finishes first, however transparently. Scribbr mitigates that last one about as well as a vendor can — publishing method and per-category failures — but an interested party's benchmark is corroborating evidence, not independent proof. We'd apply the same sentence to any self-test we ever publish, which is why HumanFlow publishes no accuracy percentage for its own AI detector without published methodology — and why that tool doesn't promise to out-judge anyone, because nobody can honestly promise that.

One more honesty wrinkle, offered gently: Scribbr's FAQ pages cite "68% in the best free tool" while the revised study table shows free tools at 78% — and as of our August 2026 check the live page had moved again, to "84% in a premium tool or 68% in the best free tool". The number moved 10 points without an announcement, which is the reminder: even the transparent vendor's figures need a date attached.

What the published record says about accuracy — beyond Scribbr's own test

Put Scribbr's self-reported 78–84% next to the category's independent evidence and it looks less like modesty and more like realism. The largest peer-reviewed comparison, Weber-Wulff et al., put all fourteen tools it tested below 80% accuracy — only five cleared 70% — and measured 26% on machine-paraphrased text. OpenAI retired its own classifier in July 2023 after it caught just 26% of AI text while wrongly flagging 9% of human writing. And Liang et al. (Patterns, 2023) found seven detectors falsely flagged an average of 61.22% of human-written TOEFL essays by non-native English speakers — near-perfect on US 8th-graders' essays, disastrous on ESL writers. Scribbr's detector wasn't in that study, but it uses the same statistical approach the study indicts, and Scribbr's core audience includes exactly the international students the bias falls on. If you're one of them, our guide to detector false positives explains who gets misflagged and why.

The arithmetic that should sit under every score: suppose a detector is 99% specific — better than most evidence supports — and a university screens 10,000 honest papers. That's roughly 100 innocent students flagged. If one submission in ten is actually AI-written and the tool catches 95% of those, you get about 950 true catches and, at scale across all submissions, roughly one false accusation for every ten true ones. Scribbr's paragraph-level categories and cautious language are a better fit for that reality than a confident single percentage. A tool built for self-checking can afford honesty; a tool sold for enforcement has to pretend.

What we have not done

We have not run our own hands-on test of Scribbr's detector. Everything above rests on vendor claims and published research, and that is exactly how you should weight it. We would rather say so than publish a table of numbers we did not measure — and it is why the next section hands you the method instead of asking you to trust ours.

Who it serves, and who should look elsewhere

The fit is best for exactly who Scribbr built it for. A student self-checking before submission gets unlimited free scans, no account, a privacy promise, and mixed-content categories that map onto how students actually use AI — that's the strongest free-tool user experience of the three detectors in this review cluster. An ESL student worried about bias gets a vendor that at least publishes its limits, though no statistical detector, Scribbr's included, removes the underlying risk. A thesis writer already inside Scribbr's proofreading and citation ecosystem gets detection alongside plagiarism checking in one familiar place.

Look elsewhere if you need current-frontier model coverage stated in writing (the page's named models lag), scanning beyond 1,200 words without slicing, an API or bulk workflow (content agencies should see the enterprise-oriented tools in our accuracy hub), or sentence-level rather than paragraph-level granularity — Grammarly's Pro tier and QuillBot both go finer-grained. And if you're an instructor: Scribbr's own research is your reason not to treat any score as proof. The vendor told you its best tool misses 40% of the hardest category. Believe it.

Limits, stated plainly

The 1,200-word cap means chapter-length work gets scanned in pieces, and stitching per-piece verdicts into a document-level judgment is on you. Paragraph-level output is coarser than the sentence-level highlighting elsewhere. Stated model coverage is a generation behind the frontier until verified otherwise. Four languages trail QuillBot's twenty-plus. The self-published study, admirable as it is, is small, partly stale on generators, and authored by an interested party. And the detector shares the category's foundational limit: it measures how machine-typical text is, not who wrote it, so its errors concentrate on honest writers whose prose happens to look statistical — formulaic academic structure most of all, which is an awkward fact for a tool aimed at academia.

None of that reverses the verdict. As free detectors go, Scribbr's is well-scoped, honestly framed, and unusually well-documented by its own maker. In this category, that combination is nearly unique.

FAQ

Is Scribbr's AI detector really free? Yes — up to 1,200 words per scan, unlimited scans, no sign-up, and Scribbr states submissions aren't stored or shared. A premium version with higher claimed accuracy and GPT-4 coverage exists; its current pricing wasn't listed on the pages we fetched, so check the live site.

How accurate is Scribbr's AI detector? By Scribbr's own published research (last revised July 2026): 84% for premium and 78% for free under its 30-text methodology, with only 60% of paraphrased or mixed texts caught even by premium. Those self-reported numbers are more modest — and more believable — than the category's typical 99% claims, but the sample is small and partly based on older generators.

Can Scribbr detect ChatGPT and newer models? Scribbr's pages state the free tool detects GPT-2, GPT-3, and GPT-3.5 with "average accuracy," and premium adds GPT-4 with "high accuracy." Coverage of newer frontier models isn't clearly claimed on the fetched pages, so treat current-generation detection as unverified.

What makes Scribbr different from other detector vendors? It publishes comparative research grading twelve detectors including its own, with methodology, and reports results where its free tool ties a competitor's and its premium tool visibly struggles on mixed text. Almost no other vendor volunteers numbers like that about itself.

Is Scribbr's detector safe for my unpublished thesis? Scribbr states it doesn't store or share submissions, which is the right promise for academic work. Like any vendor privacy claim it rests on trust rather than user-side verification, so institutions with strict data rules should confirm terms before pasting embargoed research.

Can a teacher use a Scribbr score as proof of AI use? No — and Scribbr agrees, writing that detectors "shouldn't be treated as absolute proof." Its own best tool missed 40% of paraphrased/mixed texts in its own test, and the category's documented false-positive bias hits non-native English writers hardest. Use scores to start conversations, alongside drafts and process evidence.

Is Scribbr's detector better than QuillBot's or Grammarly's? Scribbr's own test scores its free tool level with QuillBot's (78% each). Practically: Scribbr wins on unlimited free scans, no sign-up, and published research; QuillBot on languages and sentence-level detail; Grammarly on ecosystem reach and cautious framing. All three share statistical detection's structural limits.

Key facts

  • Scribbr's free detector: 1,200 words per scan, unlimited scans, no sign-up, with a stated no-storage privacy policy (scribbr.com/ai-detector, fetched August 2026).
  • Scribbr's own comparative study (last revised July 21, 2026) scores Scribbr premium 84%, Scribbr free and QuillBot 78%, Originality.ai 76%, GPTZero 52%, OpenAI's retired classifier 38% — across 30 texts in six categories.
  • Even Scribbr's top-scoring premium tool caught only 60% of paraphrased or mixed human-AI texts in its own test.
  • Stated free-tier model coverage: GPT-2/3/3.5 ("average accuracy"); premium adds GPT-4 ("high accuracy"); no newer model is named for either tier.
  • Languages: English, Spanish, German, French; results are paragraph-level across four origin categories.
  • Liang et al. (Patterns, 2023): 61.22% average false-positive rate on non-native English writers' essays across seven detectors — Scribbr wasn't tested, but shares the statistical approach.
  • OpenAI retired its own classifier in July 2023 at 26% detection and 9% false positives — the tool that finished joint-last in Scribbr's comparison.

Sources

  1. Scribbr — AI Detector page (scribbr.com/ai-detector), fetched August 2026.
  2. Scribbr — "Best AI Detectors" comparative research (scribbr.com/ai-tools/best-ai-detector/), last revised July 21, 2026, fetched August 2026.
  3. Scribbr — FAQ, "How accurate is Scribbr's AI detection software?" (scribbr.com/frequently-asked-questions/), fetched August 2026.
  4. Liang, W. et al., "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023.
  5. OpenAI — AI text classifier retirement announcement, July 2023.
  6. Weber-Wulff, D. et al. (2023). Testing of detection tools for AI-generated text. International Journal for Educational Integrity 19:26.
All postsPublished by The HumanFlow team