First, the thing to do today
Before anything else: preserve your draft history. Do not tidy the document, do not delete old versions, do not start a clean copy, and do not rewrite the submitted text. If you wrote this essay, the record of you writing it already exists, and it is worth more than any argument you can make about detector accuracy.
- Google Docs — File → Version history → See version history. It keeps a fine-grained record you can scrub through. Note that browsing earlier versions requires edit permission on the file, so make sure you still hold it.
- Microsoft Word — AutoSave revisions in OneDrive or SharePoint, plus any emailed drafts.
- Everything around the document — notes, outlines, photographs of handwritten planning, library loan records, browser history for the sources you read, messages where you talked about the assignment.
Timestamps spread over days are the point. They describe a process. A detection score describes a finished artefact and cannot speak to how it came to exist.
What the number actually is
A detector never looks for AI. It measures two surface properties — how predictable each word is given the words before it, and how much your sentence lengths vary — and converts them into a percentage. Language models produce predictable words and even rhythm, so they score one way. So does a careful writer working in a second language. So does anyone taught to write a five-paragraph essay properly. The score cannot tell those apart, because it is not measuring authorship. Our breakdown of the mechanism goes through it in full.
This is not a fringe criticism. In the largest peer-reviewed test of these tools, Weber-Wulff and colleagues examined fourteen detectors and concluded they “are neither accurate nor reliable” — every one scored below 80% accuracy, and on machine-paraphrased text overall accuracy was 26%. Perkins and colleagues measured seven detectors at 39.5% accuracy on unaltered AI text, dropping to 22.2% once the text had been lightly altered.
And the bias is documented and specific. Liang and colleagues found that seven detectors falsely flagged 61.22% of human-written TOEFL essays while judging essays by native- speaking US eighth-graders almost perfectly. If English is not your first language, this is not bad luck. It is a known property of the method, and you may say so.
The vendor’s own numbers help you
Turnitin claims 98% accuracy — but only for documents where more than 20% of the text is flagged, and it hides scores between 1% and 19% behind an asterisk because its own validation showed that range is unreliable. If your flagged percentage is low, the company that built the tool has already said the number should not be printed, let alone acted on.
Turnitin’s chief product officer, Annie Chechitelli, also told BestColleges in April 2023 that the system is deliberately tuned to miss: “we are estimating that we find about 85% of it. We let probably 15% go by in order to reduce our false positives to less than 1 percent.” That cuts both ways in a hearing, and the half that matters to you is this — the vendor designed the tool around the knowledge that false accusations are the worse error.
You are not the first, and institutions have noticed
Vanderbilt disabled Turnitin’s AI detector in August 2023 and published the arithmetic: 75,000 papers submitted in 2022, and at Turnitin’s own claimed error rate, “around 750 student papers could have been incorrectly labeled.” Its conclusion was that it did “not believe that AI detection software is an effective tool that should be used.”
Since then: Yale, Georgetown, the University of Pittsburgh, Johns Hopkins, the University of Alabama and Curtin University have all turned the feature off. Washington State University cancelled its Turnitin AI detection contract outright in February 2026, and its provost’s memo disclosed the reason plainly: between 2023 and 2025, a third of its academic-integrity hearings involving AI allegations ended in a finding of not responsible, because a detector score had been submitted as the only evidence.
Others never adopted it. UC Berkeley ran a pilot and opted out. Syracuse declined to license it. NYU’s provost’s office states that it does not “believe any current AI detectors work well enough to recommend their use.” Cambridge “does not encourage the use of AI detection software given their proven inaccuracies and unreliability.” Monash has approved no detector at all.
Ask which camp your institution is in, and ask for the policy in writing. If your department is running a tool the university has not approved, that is relevant to your case.
What to say, and what not to say
- Ask for the specific allegation and the evidence behind it. A percentage is not an allegation. What passage, and what supports the claim beyond the score?
- Offer your process, not your protest. Draft history, notes, timeline. Offer to talk through the argument of the essay — someone who wrote it can.
- Ask whether the score alone is sufficient under policy. Increasingly it is not, and asking makes the standard explicit.
- Do not admit to something you did not do to end the meeting. This is the single most costly mistake, and it is common, because these meetings are frightening and an admission feels like the fastest way out. It closes the appeal routes your evidence would have won.
- Do not alter the submitted text. Rewriting or running the work through a rewriting tool after an accusation reads as tampering, whatever your intention.
- Ask what support you are entitled to. Most institutions have a student union, advocate or ombudsperson who does this regularly. Use them.
Where you are studying changes the route
The evidence above travels anywhere. The escalation route does not — it is set by national rules, and it is the part most guidance written for a US audience leaves out.
- England and Wales: once your university's internal process ends you can take the case to an independent body, on a clock that starts with a specific letter. The appeal process, and the 12-month deadline.
- The UK generally: no regulator requires AI detection at all, which changes what a panel can claim it is obliged to do — who regulates this in the UK.
- India: the UGC regulation your department is citing grades similarity, and AI writing produces almost none — what the UGC rules actually say.
Where we stand, since we sell one of these tools
HumanFlow makes a humanizer and a detector, so treat this section as an interested party speaking. We will not tell you that our detector — or anyone’s — can prove who wrote a document. It cannot. No detector can, and the peer-reviewed record above is the reason we do not make that claim anywhere on this site.
What a detector is genuinely useful for is seeing your own writing the way an institution’s tool might see it, before you submit. That is a different job from proving authorship, and it is the only one we think the technology honestly does. If you are already under investigation, the useful artefact is your draft history, not another score.
Before arguing about the number
Establish what it measures. An asterisk is Turnitin withholding a figure it does not trust, 20% is where it starts displaying one rather than a limit, and a GPTZero probability is confidence in its own verdict rather than a share of your document — what a score actually means.
What the tool will not read
A missing score is often a file limit rather than a verdict. Turnitin’s AI report needs at least 300 words of prose and does not accept slides at all — see PDFs and PowerPoint.
Related reading
- Why AI detectors flag human writing — the mechanism, and who it hits hardest.
- How accurate AI detectors actually are — every published figure, with the conditions attached.
- Turnitin’s AI detector, reviewed — what it does, what it claims, and where the claims stop.
- Turnitin false positives: how often innocent writing gets flagged