humanflow
AI detection · The HumanFlow team · 13 min read

Can AI Detectors Detect Edited AI Text? The Mixed-Authorship Problem Nobody Has Solved

Lightly edited AI text usually still gets flagged. Deeply rewritten text usually doesn't — because its authorship genuinely changed. Blends are where detectors fail.

It depends on how deep the editing goes — and that dependence is the whole story. Lightly edited AI text usually still gets flagged, because a few word swaps barely move the statistics. Deeply rewritten text usually doesn't, because its statistics genuinely become human. In between sits the blend: part-AI, part-human documents that detectors handle worst, by their makers' own admission.

Pure cases are easy. A pasted-in ChatGPT essay and an essay typed from scratch sit at opposite ends of a statistical spectrum, and detectors separate those ends with the 90-plus-percent accuracy their marketing celebrates. But almost nobody who uses AI seriously lives at either end. Real documents are drafted by a model and restructured by a person, or drafted by a person and smoothed by a model, or assembled from both in alternating paragraphs. That middle territory is where detection actually gets used — and it's where detection is weakest. This might be the most consequential fact in the entire AI-detection debate, so let's take it apart properly.

The documented case: Turnitin meets a blended draft

The clearest public demonstration remains the Washington Post's test of Turnitin in April 2023. Reporter Geoffrey Fowler ran real student writing and AI-assisted writing through Turnitin's then-new detector. The results made news for two failures, not one. The detector flagged writing by an innocent student — a straight false positive. And it struggled badly with mixed human/AI drafts, misjudging how much of the blended documents was machine-written. Turnitin's response was notable for its honesty: the company acknowledged that blended documents are the hard case for its technology.

Sit with that acknowledgment. The vendor with the largest deployment in education — over 200 million papers screened in its first year — conceded in its launch month that the most realistic scenario, a student mixing their own words with a model's, is precisely where its accuracy claims get shaky. Turnitin's fine print says the same thing structurally: its headline 98% accuracy and sub-1% false-positive figures apply only to documents where more than 20% of the text is flagged, and scores from 1–19% display as an asterisk rather than a number because the company itself doesn't trust low-range readings enough to print them. An asterisk is what a mixed document often earns. The asterisk is the mixed-authorship problem wearing corporate typography.

What editing does to the statistics, step by step

To see why blends break detectors, you need only the two quantities detectors measure. Perplexity: how predictable each next word is. Burstiness: how much sentences vary in length and shape. Machine text is smooth and even; human text is spiky and irregular. A classifier slices your document into segments, scores each one, and aggregates. (Full mechanics at our accuracy hub and the explainers linked from it.)

Now watch what happens as a human edits an AI draft, in ascending order of depth.

Fixing typos and swapping a few words changes almost nothing. The sentence skeletons — the part carrying most of the statistical signal — remain model-built. Perplexity ticks up a hair because your synonym is slightly less probable than the model's choice. The segment still scores machine-typical. Detectors catch this level of editing routinely, and every credible test agrees.

Rewriting sentences here and there starts to matter, unevenly. Each rewritten sentence becomes a human-typical data point sitting inside machine-typical surroundings. Document-level scores start to wobble, because the aggregate now averages two populations. This is where different tools begin disagreeing with each other about the same text — one vendor's threshold calls the mix AI, another's calls it human, and both are defensibly reading the same ambiguous statistics.

Restructuring and rewriting most of the text in your own words flips the signal. Your argument order, your transitions, your vocabulary, your uneven rhythm — at this depth the majority of segments score human-typical because they are human-typical. Whatever machine scaffolding survives is diluted below most thresholds. The document reads, statistically, as yours.

The academic evidence backs the trajectory and adds a warning about instability. The RAID benchmark (ACL 2024) — six million generations, twelve detectors, eleven adversarial attack types — found paraphrase attacks produced wildly inconsistent effects across detectors. Its own table makes the point better than any summary: run the same paraphrase attack against the four commercial tools it tested and Originality's accuracy rose, from 85.0 to 96.7, while Winston's fell from 71.0 to 52.6 and ZeroGPT's from 65.5 to 46.7. GPTZero barely moved, down 2.5 points. Same attack, same corpus, same day — a 30-point spread in which direction the tools even went. Perkins et al. put a number on the cost of editing that nobody should have to squint at: seven detectors averaging 39.5% accuracy before any adversarial technique was applied, and 22.2% after. That is the peer-reviewed measurement of what editing does, and it is far more damning than any range. The pattern is not "editing defeats detectors." The pattern is "editing pushes detectors into the zone where they disagree with each other and with the truth, in both directions."

Both directions matters. The same statistical logic that lets deeply edited AI text score human also lets heavily self-edited human text score AI. Writers who revise toward uniformity — polishing every sentence to the same smooth length, sanding off their own irregularities — manufacture exactly the low-burstiness signal detectors flag. That's one reason meticulous self-editors appear on every documented false-positive risk list alongside non-native English speakers, and it's covered in depth in our false positives explainer.

The spectrum, made concrete

Editing depthWhat actually happenedWhat the statistics showWhat detectors typically doWhose writing is it, honestly?
None — raw pasteModel wrote everythingUniformly machine-typicalFlag it, reliably (90%+ in most tests)The model's
Light touch-upTypos fixed, a few words swappedBarely changedStill flag it, usuallyThe model's, with your fingerprints
Partial rewriteSome sentences yours, skeleton still the model'sMixed segments; unstable aggregateDisagree with each other; scores land mid-range or asterisk territoryGenuinely shared — and contested
Deep rewriteYour structure, your words; AI ideas survive as raw materialMostly human-typicalUsually score it humanYours, built on assisted research/drafting
AI as editorYou wrote it; model smoothed grammarMostly human-typical, uniformity nudged upUsually human — with false-positive risk if smoothing is heavyYours

Read the last column against the third one. They track each other. That correlation is the most underappreciated fact in this debate, and it's where this article has been heading.

Why sentence-level readouts matter more than the big number

A single document-level percentage is the wrong instrument for a mixed document, almost by definition. "34% AI" stapled to a blended essay answers nothing anyone actually needs to know: which 34%? The literature review the student openly drafted with permitted AI help, or the analysis section they claim as their own? One number flattens the only question that matters in a mixed-authorship dispute — where, specifically, the machine-typical text sits.

Sentence- or segment-level readouts at least point at something checkable. If a tool highlights specific passages as machine-typical, a writer can respond to the specific passages — with drafts, version history, or an honest account of their process — and a reviewer can weigh the highlighted text against what they know of the writer's voice. The conversation becomes about evidence instead of about a percentage neither party can interrogate. This is why our own AI detector reports sentence-level results rather than a single verdict, and why we publish no accuracy percentage for it without published methodology. It does not promise to beat, predict, or out-judge any other tool — nobody can honestly promise that, least of all on mixed documents, which we've just spent an article explaining are the hard case for the entire category.

The same logic should discipline how anyone reads a detector report. A mid-range score on a blend is not "34% cheating." It's a statistical estimate, from a tool its own vendors admit is weakest on blends, of how much of the text is machine-typical — which is not the same as machine-written, and says nothing about whether the machine's contribution was permitted. Every one of those gaps has to be closed by a human before a score becomes an accusation.

The legitimacy question, faced squarely

Here's the reframe this whole topic needs. When someone deeply edits AI output — restructures it, rewrites it in their own voice, takes ownership of every claim — and a detector then scores it human, the detector has not been beaten. It has returned a defensible reading of a document whose authorship genuinely changed. Detectors estimate a fact about text: how machine-typical it is. Deep editing changes that fact. The score moved because the truth it estimates moved. Calling that "evading detection" makes exactly as much sense as saying a repainted house "evaded" the paint inspector.

This is not a loophole argument, and the difference is worth stating with no hedging. If your course, employer, or publication prohibits AI assistance, then AI-assisted work violates the rules no matter how deeply you edit, no matter what any detector says — disguising prohibited assistance is a conduct problem, not a statistics problem, and a human-looking score doesn't launder it. The legitimacy of deep editing depends entirely on whether the assistance itself was permitted, and on whether you're honest about it when asked. Where AI use is allowed — as it increasingly is, with disclosure, in workplaces and a growing share of classrooms — the writer who transforms a model's draft into their own considered prose is doing the thing writing teachers have always described as writing: revision until the words are yours.

What deep editing is not is a shortcut. Genuinely rewriting a 2,000-word draft in your own voice — checking its claims, restructuring its argument, replacing its phrasing — takes real hours and real understanding; anyone who's done it knows it can approach the effort of drafting from scratch. That's precisely why it changes both the statistics and the authorship: the labor is the authorship. Tools can support that process — our AI humanizer is built for voice-and-tone revision, and it carries no promise of beating any detector, because that isn't what revision is for and no such promise can be honestly made. The work either becomes yours or it doesn't. Software can assist the becoming; it can't substitute for it.

And what light editing is not is transformation. Swapping synonyms into a pasted essay changes neither the statistics much (detectors mostly still catch it) nor the authorship at all (the thinking is still the model's). The people most likely to be caught by detectors are those trying to buy the deep-edit outcome at the light-edit price. The statistics, for once, are on the side of the effort.

Where this sits in the bigger detection picture

Mixed authorship is the third leg of a triangle this pillar has been assembling. The per-model question — can detectors catch Claude, or GPT, or Gemini — turns out to matter little: detectors flag machine-typical statistics regardless of brand. The freshness question — can detectors catch models newer than their training data — matters for a few weeks after each launch, then vendors retrain and the gap closes. This question is the one that doesn't resolve. Model brands converge; launch gaps close; but the spectrum between "the model wrote it" and "I wrote it" is permanent, because it's not a gap in the technology. It's ambiguity in the thing being measured. A blended document does not have a clean binary answer for a classifier to find, and no amount of retraining conjures one.

Which is why every serious accuracy conversation ends at the same place: a detector score is an estimate of a statistical property, produced by a tool that is strong at the ends of the spectrum and weak in the middle, where most real contested cases live. Used as one signal among several, read at sentence level, weighed by a human who understands its failure modes — useful. Used as a verdict on a blend — indefensible, by the vendors' own fine print.

The pure cases were always easy. The blends are where the people are.

FAQ

Can AI detectors detect AI text after light editing? Usually yes. Fixing typos and swapping scattered words leaves the model-built sentence structures intact, and those carry most of the statistical signal detectors read. Credible tests consistently show lightly edited AI text still gets flagged at high rates.

Does paraphrasing AI text make it undetectable? No — it makes results unstable, which is different. The RAID benchmark (ACL 2024) ran one paraphrase attack against four commercial detectors and sent them in opposite directions: Originality's accuracy rose from 85.0 to 96.7, while Winston's fell from 71.0 to 52.6 and ZeroGPT's from 65.5 to 46.7. Some vendors have extended detection toward this ground: Turnitin added paraphrase detection in July 2024 and bypasser detection in August 2025, publishing no accuracy figure for either, and limiting the bypasser feature to English. Unstable is not safe.

How much editing does it take before AI text scores as human? There's no published bright line, and it varies by tool and threshold. The documented pattern: word-level changes move little; sentence-level rewriting creates mixed, tool-dependent results; restructuring plus rewriting most of the text in your own words typically crosses into human-typical territory — because at that depth the text largely is human-written.

What did the Washington Post test actually show? In April 2023, the Post ran student and AI-assisted writing through Turnitin's new detector. It flagged an innocent student's work and misjudged mixed human/AI drafts, and Turnitin acknowledged that blended documents are the hard case for its technology. It remains the clearest public demonstration of the mixed-authorship problem.

Why does Turnitin show an asterisk instead of a score below 20%? Because Turnitin itself doesn't consider 1–19% readings reliable enough to display as numbers. Its 98% accuracy and sub-1% false-positive claims apply only above the 20% line. Mixed documents frequently land in exactly that asterisk zone — the vendor's own interface conceding the blend problem.

Can a detector tell whether AI editing of human writing counts as AI text? Not really — it can only report that text is statistically smooth. A human draft heavily polished by a model may drift toward machine-typical scores, and a meticulous human self-editor can produce the same signal with no AI at all. That ambiguity is a core false-positive risk, not a solved problem.

Is deeply editing AI output cheating? It depends entirely on your context's rules, not on the detector outcome. Where AI assistance is prohibited, deep editing of AI output still violates the policy — full stop. Where assistance is permitted, transforming a draft into your own voice, with honest disclosure where required, is revision — the score reflecting that isn't evasion, it's measurement.

Should institutions rely on document-level scores for mixed submissions? No. Vendors' own conditions exclude the blend zone from their accuracy claims, and a single percentage can't say which passages are machine-typical or whether any assistance was even prohibited. Sentence-level readouts plus human review of drafts and process are the defensible minimum.

Key facts

  • Washington Post, April 2023: Turnitin's detector flagged an innocent student's writing and struggled with mixed human/AI drafts; Turnitin acknowledged blended documents are its hard case.
  • Turnitin's fine print: the 98% accuracy / <1% false-positive claims apply only when more than 20% of a document is flagged; scores of 1–19% display as an asterisk, not a number — the low range vendors themselves won't print.
  • RAID benchmark (ACL 2024): across 12 detectors and 6M+ generations, paraphrase attacks swung accuracy from +16.2 to −15.4 points depending on the detector — modified text is where tools disagree most.
  • Perkins et al. (IJETHE 21:53, 2024): seven detectors at 39.5% baseline accuracy, falling to 22.2% once adversarial editing was applied.
  • Scale of the stakes: Turnitin screened 200M+ papers in its first year (April 2023–April 2024); ~11% showed 20%+ AI writing — millions of documents, many of them inevitably blends.
  • Heavy human self-editing toward uniform prose produces the same low-burstiness signal as AI text — meticulous self-editors appear on documented false-positive risk lists alongside non-native English writers.

Sources

  1. Fowler, G., "We tested a new ChatGPT-detector for teachers. It flagged an innocent student." The Washington Post, April 2023.
  2. Turnitin, AI writing detection FAQ / transparency documentation (accuracy claim conditions, asterisk range, minimum length).
  3. Dugan et al., "RAID: A Shared Benchmark for Robust Evaluation of Machine-Generated Text Detectors," ACL 2024 (aclanthology.org/2024.acl-long.674).
  4. Turnitin, first-anniversary data release, April 2024 (papers screened, AI-share distribution).
  5. Liang et al., "GPT detectors are biased against non-native English writers," Patterns (Cell Press), 2023 — for the false-positive mechanism on low-variance human prose.
All postsPublished by The HumanFlow team