Is a low perplexity score proof of AI writing?
No. A low perplexity score establishes that your writing was predictable, which is a property of the prose rather than a fact about its author.
Last reviewed 15 August 2026 · The HumanFlow team
Perplexity measures how surprised a language model is by each word given the words before it. Model output scores low because models select likely words by construction. The inference detectors make — low perplexity, therefore machine — runs the implication backwards, and it only holds if human writing is reliably unpredictable.
It is not. Formulaic writing is predictable because the genre requires it: a methods section, a lease, a lab report, a set of assembly instructions. Careful writing in a second language is predictable because sticking to constructions you are confident in is what competence under pressure looks like. Both are human, and both score the way machine text scores.
The measured consequence is not subtle. In peer-reviewed testing published in Patterns, seven detectors flagged 61.22% of TOEFL essays written by non-native English speakers as AI-generated — 89 of 91 essays tripped at least one tool — while classifying essays by native-speaking US 8th-graders almost perfectly. The same property that produced those flags is the one perplexity measures.
So a low score is a reason to look at the writing, and it is not evidence of authorship. It cannot distinguish a machine from a person writing plainly, because there is nothing in the measurement that could.
When this answer changes
If watermarking becomes widespread, the picture changes for text from participating models — a watermark records what actually happened at generation time rather than inferring it from style. It still says nothing about text carrying no watermark, which includes anything from a model run locally.
It would also change if a detector published a false positive rate measured on the specific kind of writing you produce, at the threshold your institution uses. No vendor currently does.
Where to go next
- What perplexity actually measures — The mechanism, with two sentences carrying identical information and a note on which words move the score.
- Why AI detectors flag human writing — The wider failure, including who it lands on hardest.
- How to prove you wrote it — If this has already happened, the evidence that actually helps.
- Who this lands on hardest — The same mechanism, measured on a population, with the study that reported 61%.
Sources
One of our direct answers.