Most AI detectors need at least 100–300 words before their scores mean anything, and several refuse to score below a hard floor — Turnitin requires roughly 300 words, Winston AI 500 characters, Copyleaks 255 characters, Originality.ai 100 words. Below those minimums, a score isn't a weak signal. It's noise wearing a percent sign.
That single fact should change how you treat detector results on emails, discussion-board posts, tweets, and short-answer questions — which is to say, most of the text people actually paste into these tools. A teacher who runs a student's two-sentence answer through a free checker and gets "87% AI" hasn't learned anything about the student. They've learned something about the checker: it was willing to print a confident number on a sample too small to support any number at all.
This post explains why length is the load-bearing variable in AI detection — in plain language, no statistics degree required — then lists the documented minimums across major tools, and finishes with practical guidance for the short-text situations where detectors are most often misused. It pairs with two companion pieces: why detectors disagree with each other and how vendors choose their thresholds; short text is where both of those problems get worse at once.
Brief disclosure: we build a detector and humanizer ourselves at HumanFlow, so read our vendor commentary with that in mind. Our policy is to quote vendors precisely, with their fine print, and mark anything we couldn't verify.
Why statistical detection needs hundreds of words
AI detectors don't recognize AI the way you recognize a friend's face. They measure statistical properties of text — chiefly perplexity (how predictable each next word is to a language model) and burstiness (how much sentence length and structure vary) — and compare those measurements against patterns typical of machine output. The full mechanism is covered in how AI detectors work; what matters here is that every one of those measurements is an average across the text. And averages computed from small samples are wildly unstable.
The coin-flip analogy gets you 90% of the way. Flip a fair coin 1,000 times and you'll get close to 500 heads; if you instead got 700, you could be quite sure the coin was rigged. Flip it six times and get four heads — 67%! — and you've learned essentially nothing, because fair coins produce four-of-six constantly. Same coin, same test, same arithmetic. The only difference is sample size, and sample size is the difference between evidence and coincidence.
Words are the detector's coin flips. Every word in your text is one small observation: predictable or surprising, typical or odd. A 1,500-word essay gives the detector 1,500 observations — enough for genuine patterns to emerge from the noise, the way 1,000 flips expose a rigged coin. A 40-word email gives it 40 observations. At that size, perfectly human writing routinely "looks" machine-like by chance, and machine writing routinely looks human, for the same reason six coin flips routinely look rigged.
There's a compounding problem: short text isn't just a small sample, it's a biased one. The shorter the message, the more it's built from formulas — greetings, sign-offs, stock transitions, template sentences. "Thank you for reaching out. I'll get back to you by Friday." is maximally predictable, low-variance prose, which is precisely the statistical fingerprint detectors associate with AI. Short human writing converges on machine-typical patterns by nature, not by chance. So below a few hundred words, detectors face small-sample noise layered on top of genuinely machine-looking human conventions. That's why the failure isn't fixable with a better model — you can't average your way out of not having enough to average. A vendor's threshold dial can't rescue it either: as our thresholds piece explains, thresholds only trade one error for another, and on short text both error rates balloon simultaneously.
Vendors know all this, which brings us to the floors.
Documented minimums across major tools
We fetched or verified each vendor's stated minimum during writing (2026). Where a vendor publishes no hard floor, we've said so rather than inventing one. Note the units — some vendors count words, others characters (roughly 5–6 characters per English word, so 255 characters is in the neighborhood of 45–50 words).
| Tool | Stated minimum | Notes (vendor's own framing) |
|---|---|---|
| Turnitin | ~300 words of continuous prose | Won't generate a score below this; built primarily for long-form English prose, not lists/code (Turnitin FAQ) |
| Winston AI | 500 characters (≈80–100 words) | Hard floor for pasted text in the editor (Winston help center) |
| Copyleaks | 255 characters web platform; 350 via browser extension | "Our models need a certain volume of text to determine AI presence accurately" (Copyleaks site) |
| Originality.ai | 100 words (website scans); no API minimum | Its own help docs say accuracy diminishes below 100 words even via API |
| Sapling | No hard floor; "much more accurate after 50 or so words" | Accuracy claims explicitly conditioned on longer texts (Sapling site) |
| GPTZero | No published hard minimum | FAQ: accuracy rises with more text; document-level scores more reliable than paragraph; paragraph more than sentence (GPTZero FAQ) |
| Scribbr | No published minimum; 1,200-word max per free submission | Recommends "scanning longer pieces of text rather than individual sentences or paragraphs" (Scribbr site) |
Two honest observations about that table.
First, the floors vary by a factor of six — from Sapling's soft ~50 words to Turnitin's ~300. That's not because one vendor knows something the others don't. It's the threshold story again: a floor is itself a business decision about how much noise you're willing to print. Turnitin, whose scores can trigger misconduct hearings, sets the highest bar and refuses to score below it — the same conservative philosophy behind its asterisk range. Consumer tools, where a wrong answer costs nobody a degree, let you scan a couple of sentences and get something. Getting something is not the same as learning something.
Second, a minimum is a floor, not a safety line. Crossing 255 characters doesn't switch a detector from "noise" to "reliable" — reliability climbs gradually with length, and every vendor above says so in its own words. GPTZero's document-beats-paragraph-beats-sentence ranking is the cleanest vendor admission that the same model gets less trustworthy as the sample shrinks. Treat vendor minimums as "below this, we won't even pretend," not "above this, trust us."
What this means for the text people actually check
Now apply the table to real life, because the collision is constant: an enormous share of everyday writing sits under or near every floor on that list.
Emails. A typical business email runs 50–125 words — under Turnitin's floor, around Winston's, hovering at Copyleaks'. It's also the most formulaic genre most of us write: greeting, boilerplate courtesy, sign-off. Running an email through a detector and acting on the result — as some managers reportedly do when they suspect AI-written cover letters or outreach — is exactly the six-coin-flips scenario, aggravated by the genre's built-in predictability. A "likely AI" verdict on a five-sentence email is compatible with a diligent human, a template, or ChatGPT, and the detector cannot tell you which.
Discussion-board posts. The standard "respond in 100–150 words" LMS assignment lands underneath Turnitin's minimum — Turnitin literally will not score it — yet instructors sometimes paste those posts into free web checkers that will. The free tool prints a number precisely because it has a lower floor, i.e., a higher tolerance for printing noise. The result is a grading decision made on the least reliable output in the entire detection industry.
Tweets and social posts. At 280 characters, a maxed-out tweet barely clears Copyleaks' floor and falls below Winston's. Detection verdicts on individual social posts are not meaningfully better than astrology, and no reputable vendor claims otherwise.
Short-answer exam questions. Two or three sentences — 30 to 60 words. This is the case that produces the worst real-world outcomes, so it deserves its own section.
Also worth naming: the "check each paragraph separately" habit. Some users split a long essay into chunks to localize AI use — which converts one adequately sized sample into five inadequate ones and multiplies the error rate. If your tool offers sentence- or paragraph-level highlighting within a full-document scan (several do, including ours), that's different: the model still sees the whole context. Splitting the input yourself destroys that context.
The teacher with a two-sentence answer
Picture the scenario this post exists for. A student's short-answer response reads oddly polished. The teacher pastes those two sentences — maybe 40 words — into a free detector. It returns "92% AI." Case closed?
Here's what actually happened, statistically. Forty words gave the model forty observations of a genre (the compressed, definition-shaped exam answer) that is formulaic by design — students are taught to write short answers in flat, information-dense, predictable sentences. Low burstiness, low perplexity: the machine-typical fingerprint, produced by a human doing exactly what the rubric asked. Layer the small-sample randomness on top and the score becomes a coin toss weighted against precisely the students who write the way school trained them to. The documented false-positive risk groups — non-native English speakers first among them, per Liang et al.'s finding that seven detectors falsely flagged human TOEFL essays 61.22% of the time on average — are hit hardest, and Liang's essays were full-length; shortening the sample makes every one of those failure modes worse. If your classroom includes ESL writers, a short-text detector score is close to the least fair evidence you could consult.
The kicker: the serious education vendor agrees. Turnitin — with more incentive than anyone to sell detection everywhere — chose to make scoring impossible below ~300 words. When the market leader in academic integrity says two sentences can't be scored, a free web tool printing "92%" on the same input isn't outperforming Turnitin. It's answering a question Turnitin declined, on principle, to pretend it could answer.
None of this means the student didn't use AI. It means the detector cannot tell you, and something else has to: ask the student to explain their answer verbally, compare the response against their in-class writing, look at edit history if the platform records it. Ten minutes of ordinary teaching beats any number a 40-word scan can produce.
Practical guidance for short text
Distilled rules, by role:
If you're checking someone else's short text (teacher, editor, manager): don't. Below ~300 words, treat every detector verdict as inadmissible — the score is noise, and acting on noise creates real harm with documented bias against non-native and formula-trained writers. If the stakes matter, use process evidence: drafts, version history, oral follow-up, comparison with known writing. For borderline lengths (300–500 words), treat scores as a weak hint at most, never a finding.
If you're checking your own short text (student, job applicant, freelancer): understand that a scary score on a short sample says almost nothing about you — and a clean score protects you just as little, since the noise cuts both ways. If a gatekeeper will run your short text through a detector anyway, the rational defenses are process-based: keep drafts, write in platforms that timestamp revisions. When we're asked whether our own AI detector should be trusted on a paragraph, we say the same thing we'd say about anyone's: single short scans are unreliable, which is why ours reports sentence-level detail within longer documents and why it doesn't promise to beat any other detector — nobody can honestly promise that.
If you must scan short text anyway: aggregate before you scan. Ten discussion posts by the same author, combined into one 1,200-word document, give a detector something statistically legitimate to chew on — one 120-word post does not. (Be careful with interpretation: a flag on the aggregate can't localize which post tripped it, only that the author's combined output trends machine-typical.) And whatever tool you use, check its stated minimum first; a tool willing to score 30 words is telling you something about its standards.
For everyone: remember what a detector measures. Machine-typicality, not authorship — and short, conventional writing is machine-typical when humans do it correctly. The shorter the text, the more "AI-like" and "well-trained human" become the same measurement.
The bottom line
Length isn't a footnote in AI detection — it's a precondition. Every mechanism these tools rely on is an average, and averages need samples. The industry's own floors (Turnitin ~300 words, Winston 500 characters, Copyleaks 255, Originality 100 words) mark where vendors stop pretending, and reliability keeps climbing well past every floor. Below them, scores are coin flips dressed as forensics; near them, weak hints; comfortably above them, useful-but-fallible signals whose remaining failure modes we cover across our detector accuracy hub. If one sentence survives from this article, make it this one: a detector score on two sentences is not evidence about the writer — it's evidence about the tool's willingness to guess.
FAQ
How many words do AI detectors need to be accurate? Hard floors range from ~50 to ~300 words depending on the tool (Turnitin ~300 words; Winston 500 characters; Copyleaks 255 characters; Originality.ai 100 words), but reliability keeps improving well beyond the minimum. A practical rule: treat anything under ~300 words as unscoreable and expect meaningfully stable results only on several hundred words or more.
What is Turnitin's minimum word count for AI detection? Roughly 300 words of continuous prose — below that, Turnitin doesn't generate an AI score at all. The system is also built primarily for long-form English prose, not lists, code, or equations, per Turnitin's own documentation.
Can AI detectors check a tweet or a text message? They'll often accept one and print a number, but the result is statistical noise: a 280-character tweet contains far too few words for perplexity and burstiness averages to stabilize. No reputable vendor claims reliable single-post detection.
Why did a detector flag my short email as AI? Short emails are formulaic by nature — greetings, stock phrases, sign-offs — which makes them look statistically machine-typical even when a human wrote every word. Combine that with small-sample randomness and false flags on short text are common and meaningless.
Is a "0% AI" result on a short answer proof it's human-written? No. The noise cuts both ways: short samples produce unreliable scores in either direction, so a clean result on 50 words exonerates no one, just as a flag on 50 words convicts no one.
Should teachers run discussion posts through AI detectors? Not individually — a 100–150-word post is below Turnitin's own scoring floor, and free tools that will score it are printing noise. If checking feels necessary, aggregate multiple posts by the same student into one longer document, and prefer process evidence (drafts, oral follow-up) for any actual decision.
Do detectors work better on one long document than several short ones? Yes, and every major vendor says so — GPTZero, for example, documents that document-level scores are more reliable than paragraph-level, which beat sentence-level. Splitting text into chunks before scanning multiplies error rates; scan whole documents and use built-in sentence highlighting instead.
Does combining short texts to reach the minimum actually work? It gives the detector a legitimate sample size, which fixes the statistical problem — but interpret carefully: a flag on the combined document tells you the author's aggregate style trends machine-typical, not which individual piece (if any) was AI-written.
Key facts
- Turnitin requires roughly 300 words of continuous prose before producing an AI score, and was built/validated primarily on long-form English (Turnitin FAQ).
- Winston AI's editor requires a 500-character minimum; Copyleaks requires 255 characters on its web platform and 350 via extension (vendor documentation, fetched 2026).
- Originality.ai website scans require 100 words, and its help center states accuracy diminishes below 100 words even on the no-minimum API (Originality.ai help center, fetched 2026).
- Sapling states its detector "becomes much more accurate after 50 or so words" and conditions its 97%+ detection / <3% false-positive claims on longer texts (Sapling site, fetched 2026).
- GPTZero documents that accuracy rises with length: document-level classification beats paragraph-level, which beats sentence-level (GPTZero FAQ, fetched 2026).
- Liang et al., Patterns (2023): seven detectors falsely flagged human-written TOEFL essays 61.22% of the time on average — on full-length essays; short samples worsen every documented bias.
- Turnitin displays 1–19% scores as an asterisk, its own acknowledgment that low-confidence output shouldn't be printed as a number — the same philosophy behind length floors.
Sources
- Turnitin — AI writing detection FAQ / transparency page (~300-word minimum, English long-form scope, asterisk range).
- Winston AI — help center, "How to Use Winston AI for Text Analysis" (500-character minimum; fetched during writing, 2026).
- Copyleaks — AI Detector page (255/350-character minimums; "our models need a certain volume of text" caveat; fetched during writing, 2026).
- Originality.ai — help center, "Minimum Word Counts for Scans" (100-word site minimum; sub-100-word accuracy caveat; fetched during writing, 2026).
- Sapling — AI Content Detector page (~50-word guidance; accuracy claims and caveats; fetched during writing, 2026).
- GPTZero — FAQ (length–accuracy relationship; document/paragraph/sentence reliability ranking; fetched during writing, 2026).
- Scribbr — AI Detector page (1,200-word free limit; longer-text recommendation; fetched during writing, 2026).
- Liang, W., et al. — "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023.