AI detection
AI detectors are widely used and widely misunderstood. These pages explain what the scores actually measure, what the research says about how often they are wrong, and what to do if you have been accused on the strength of one.
We build a humanizer and a detector, so we have an obvious interest here. We have tried to write these pages so that someone who never installs our app still leaves with a correct understanding — and we say plainly where the evidence runs out.
Are AI detectors accurate?
What peer-reviewed testing found, including a 61.22% false positive rate on essays by non-native English speakers — and why OpenAI withdrew its own detector.
Read →Why human writing gets flagged
Who false positives happen to, a worked before/after example, and the steps that actually help if your own work has been questioned.
Read →How AI detectors work
Perplexity, burstiness and thresholds, explained plainly — and the structural ceiling the method runs into.
Read →The short version
- A detection score measures how predictable text is, not who wrote it.
- False positives are common and fall hardest on people writing in a second language.
- Different detectors regularly disagree about identical text, because each picks its own threshold.
- Draft history — Google Docs version history, Word AutoSave — is the strongest response to an accusation.
- No tool can promise a particular score from a particular detector. Anything claiming otherwise is selling you something.
Pages in this section last reviewed 27 July 2026. Want the tool? See what HumanFlow does.