humanflow

AI checkers for students: what a score actually means

A clean scan is not a guarantee and a flag is not proof. Detectors disagree with each other on identical text, change without notice, and misfire most often on non-native English writers. Check your work if it helps you — just know what the number can and cannot tell you.

Last reviewed 4 August 2026 · The HumanFlow team

Why two tools give you two answers

Run the same paragraph through three detectors and you will often get three different verdicts. This surprises people, and it is the single most useful thing to understand before you let a score affect your evening.

Detectors do not know who wrote anything. They estimate how statistically typical of machine writing your prose looks, largely through perplexity — how predictable each next word is — and burstiness, the variation in sentence length and structure. Each tool was trained on different text and each picks its own thresholds, so they reach different conclusions from the same input.

Your institution's detector is the only one whose opinion has consequences, and you almost certainly cannot run it yourself — Turnitin is sold to institutions rather than students. So a third-party scan is a rough signal about your writing, not a preview of your result. Anyone selling it as a preview is misleading you.

If English is not your first language

You should know the research, because it is about you specifically. A 2023 Stanford study by Liang and colleagues, published in Patterns, ran 91 TOEFL essays — all written by real people — through seven detectors. They flagged an average of 61.22% as AI-generated. Eighty-nine of the 91 were flagged by at least one tool. The same detectors read essays by US 8th-graders almost perfectly.

The variable was not whether a machine wrote the text. It was whose English was being read. If you write careful, conventional, well-structured English because you learned it as a second language, you are in the group these tools fail most often.

That is unfair and it is also useful to know in advance. Keep your drafts. It is the single most effective thing you can do, and it costs nothing.

Self-checking, honestly framed

If a scan reduces your anxiety enough to let you submit, run one. Just hold the result loosely, and use it for the thing it is genuinely good at.

A sentence-level readout is more useful than a document percentage, because it points at specific lines. When a detector flags a sentence, that sentence is usually flat — even length, predictable phrasing, a connective doing no work. Rewriting those lines in your own voice tends to improve the essay whatever any detector would have said afterwards. That is the real value, and it survives even if the score is wrong.

What a scan cannot do is clear you. A low score is not a certificate, and chasing one is a poor use of the hours before a deadline. If you wrote the work, your best protection is your draft history, not a number from a third party.

Keep your drafts — it is the whole defence

If you take one thing from this page, take this. Google Docs and Microsoft Word both keep revision history automatically, and that history is the strongest evidence you can produce that the work is yours.

It shows the essay accumulating: the paragraph you wrote and deleted, the section you moved, the argument that changed shape in the third draft. No detector score competes with that, and no one who did not write the essay can produce it afterwards.

Do not compose in a tool that overwrites silently, and do not paste a finished piece from elsewhere into a blank document as your only record — that produces a history showing one large paste and nothing before it, which is the pattern you least want to have to explain.

If you are wrongly flagged

  1. Do not confess to end the discomfort. A surprising number of students do, believing it will resolve things faster. It converts a defensible position into an admitted violation.
  2. Ask for the actual report. Including whether the score cleared the tool's own reliability threshold. Turnitin shows results between 0 and 20% as an asterisk because its own testing found more false positives there.
  3. Gather process evidence. Version history in Google Docs or Word, notes, outlines, earlier drafts, research browser history, and previous graded work in your voice.
  4. Bring the research. The Liang study is peer-reviewed and directly relevant if you are a non-native English speaker. Institutions take published evidence more seriously than assertion.
  5. Ask what the finding rests on besides the score. Most policies require more, and asking the question calmly is not an accusation.

Our false positives guide walks through the meeting itself — what integrity officers look for, what persuades them, and what to avoid saying.

Disclosure, and why it protects you

Where your course permits AI with disclosure, disclose it. A sentence is usually enough: name the tool, say what you used it for, and say what you did yourself. "I used ChatGPT to generate an initial outline, which I restructured, and drafted the essay myself" is a complete disclosure.

The reason to do this is not only integrity. A disclosed, bounded use of AI is a documented account of your process, written before anyone questioned it. That is worth a great deal more in a difficult conversation than an explanation offered afterwards.

Where your course bans AI, no tool makes using it acceptable — ours included. And where the policy is unclear, ask before submitting rather than after. Instructors answer that question far more generously in advance than in retrospect.

Our detector, and its limits

HumanFlow's detector is free for 10,000 words a month, up to 1,500 in a single scan, and reports sentence by sentence rather than as one figure. It returns human, AI or mixed, with the emphasis on mixed, because most real documents are a blend.

It is subject to every limitation described on this page. It cannot tell you what your institution's detector will conclude, and HumanFlow does not promise to beat any detector — nobody honestly can. We also sell a rewriter, which is a commercial interest worth stating: our editorial policy explains how we handle writing about a market we compete in.

Run a free scan →

Common questions

Should I check my essay with an AI detector before submitting?
It can be useful as information, and it is worth being clear what kind. A scan tells you how one tool reads your prose today. It cannot tell you what your institution's detector will say, because different tools disagree on identical text and they change without notice.
Why do different AI checkers give me different scores?
Because each was trained on different data and each sets its own thresholds. The same paragraph can read as 12% AI on one tool and 89% on another. That spread is normal, and it is the clearest evidence that no single score should be treated as a fact about your writing.
I wrote it myself and got flagged. What now?
Do not panic and do not confess to something you did not do. Gather your version history, drafts, notes and outlines, and read our false positives guide, which walks through the meeting itself. False positives are documented, common, and appealable.
Does using Grammarly get you flagged for AI?
Mechanical corrections — spelling, punctuation, agreement — have no plausible effect on an AI score. Accepting many full-sentence rewrites is different in kind, because those sentences were generated by a model working from yours. The gradient runs from safe to genuinely risky, and where you sit depends on which features you used.
Can I prove I wrote my own essay?
Usually, and more easily than students expect. Google Docs and Word both keep revision history by default. Drafts, notes, outlines, browser history for your research and your record of previous graded work all count. Process evidence persuades in a way that arguing about a percentage does not.
Is it cheating to use AI for a first draft?
That depends entirely on your course policy, and the honest answer is that you should read it rather than take ours. Where AI is permitted with disclosure, say so plainly. Where it is banned, no tool — ours included — makes using it acceptable.
Will an AI humanizer stop me being flagged?
Nobody can promise that, and vendors who do are making a claim about systems they neither control nor observe. We sell a humanizer and we will not tell you it beats any detector. What rewriting can do is make flat, machine-shaped prose read more like you — which usually improves the writing regardless.

Sources

  1. 1.GPT detectors are biased against non-native English writers Liang, Yuksekgonul, Mao, Wu & Zou — Patterns (Cell Press), 2023
  2. 2.AI detection tools falsely accuse international students of cheating The Markup, 2023
  3. 3.Understanding false positives within our AI writing detection capabilities Turnitin, 2023
  4. 4.Turnitin's AI writing detection capabilities FAQs Turnitin Guides, 2026
  5. 5.OpenAI scuttles AI-written text detector over 'low rate of accuracy' TechCrunch, 2023
  6. 6.Contra generative AI detection in higher education assessments arXiv, 2023

Related

Why human writing gets flagged · Are AI detectors accurate? · Turnitin AI detection in full · Essays · What your instructor is reading