humanflow
AI detection · The HumanFlow team · 15 min read

AI Detectors and Translated Text: The Blind Spot Nobody Warns You About

Machine-translating human writing often trips AI detectors — one 14-tool study saw accuracy fall ~20 points on translated text. Here's why, and what to do.

Run human writing through Google Translate or DeepL and an AI detector will often flag it as machine-generated — because, in a literal sense, the output is machine-generated. In the largest independent test to date, detector accuracy on human-written text fell by roughly 20 percentage points once that text was machine-translated. If you write in one language and translate into English, you are carrying a false-positive risk most detector marketing never mentions.

That's the short version. The longer version matters, because translation sits at the intersection of two problems detectors handle badly: text that was produced by a neural network even though a human authored the ideas, and text written by people whose English doesn't match the training data. This post covers both directions of the problem — human writing that gets flagged after translation, and AI writing that people translate on purpose to dodge detection — along with the published evidence and what to actually do about it.

Translation is generation

Start with the mechanical fact that explains almost everything else here: modern machine translation is text generation.

DeepL, Google Translate, and the translation modes inside ChatGPT or Claude are all neural sequence models. They don't look up your sentence in a dictionary and swap words. They read your source text and then generate a new English text, token by token, choosing each word according to a probability distribution — exactly the way a chatbot generates an essay. The meaning is yours. The word choices are the model's.

AI detectors don't evaluate meaning. As we explain in how detectors work, they measure statistical properties of the surface text — chiefly perplexity (how predictable each next word is to a language model) and burstiness (how much sentence length and structure vary). Machine-translated text scores machine-typical on both, for the same reason chatbot output does: a neural network picked every word, and neural networks pick probable words. The translation engine smooths out your idioms, regularizes your sentence rhythms, and resolves every ambiguity in the statistically safest way.

So when a detector flags your translated cover letter, it isn't exactly wrong. It has correctly identified that a machine chose these particular English words. It has just answered a question nobody asked. You wanted to know "did a human author this?" The detector can only ever tell you "does this text look like machine output?" For translated text, those two questions have different answers — and no statistical detector can tell them apart. That distinction is the heart of this entire accuracy debate, and translation is where it shows up most nakedly.

The published evidence: what happens to human writing after translation

The best data we have comes from Weber-Wulff et al., "Testing of detection tools for AI-generated text," published in the International Journal for Educational Integrity in December 2023. Nine researchers tested fourteen detection tools — including Turnitin, GPTZero, ZeroGPT, Compilatio, and OpenAI's own classifier — against several categories of documents.

One category was built precisely for our question. The researchers wrote original, unpublished texts of about 10,000 characters in seven languages — Bosnian, Czech, German, Latvian, Slovak, Spanish, and Swedish — then machine-translated them into English, using DeepL for three documents and Google Translate for six. Every word of the underlying content was human. No chatbot was involved at any point.

The results:

  • On human-written English originals, the fourteen tools averaged 96% accuracy — they mostly, correctly, said "human."
  • On the machine-translated human texts, accuracy dropped by roughly 20 percentage points, and false accusations rose sharply, with some tools misclassifying translated human writing at rates the authors describe as dramatically higher.

Read that again from a student's chair. The same person, the same ideas, the same essay — written in Czech and translated with a standard tool — becomes several times more likely to be branded as AI-generated. The authors' overall conclusion was blunt: the tools tested are "neither accurate nor reliable," and no detector score should decide a misconduct case on its own.

One fairness note the study itself supports: the tools were not equally bad. Turnitin scored highest overall in the Weber-Wulff evaluation, and the translated-text failure mode varied a lot between tools. But no tool in the study, and no tool we know of since, publishes a validated false-positive rate specifically for machine-translated human writing. That silence is the blind spot.

The ESL connection: same bias, different mechanism

If this pattern sounds familiar, it should. It rhymes with the most cited finding in the entire detection literature.

In 2023, Weixin Liang and colleagues at Stanford published "GPT detectors are biased against non-native English writers" in Patterns (Cell Press). They ran 91 human-written TOEFL essays — real essays by real non-native speakers — through seven commercial and open detectors. On average, 61.22% of those essays were falsely flagged as AI-generated. Eighty-nine of the 91 were flagged by at least one detector; 18 were flagged by all seven. The same detectors were near-perfect on essays written by native-speaking US eighth graders. (Turnitin was not among the seven tested, though it uses the same statistical approach.)

The mechanism is the mirror image of the translation problem. Non-native writers tend toward more predictable vocabulary and more uniform sentence structures — they were often taught to write that way, and they are working with a smaller stock of idiomatic English. Predictable vocabulary and uniform structure are exactly what perplexity and burstiness measurements read as "machine." Liang's essays were flagged because human writing in careful learner's English statistically resembles model output. Weber-Wulff's texts were flagged because a model literally wrote the English words. Different causes; identical verdict; identical harm.

Now stack the two. An international student who drafts in her first language and translates — a workflow writing centers have recommended for decades as a legitimate way to think clearly before wrestling with English — is exposed to both failure modes at once. Her drafting language patterns pull the perplexity down, and the translation engine's word choices pull it down further. She has done nothing that any academic integrity policy prohibits, and she is close to the worst-case profile for a false flag. We've written more about who gets falsely flagged and why in our false positives guide, and the pattern holds across every population studied: the people most likely to be wrongly accused are the people least equipped to fight the accusation.

The other direction: translating AI text to dodge detectors

Flip the problem around. Some people generate an essay with ChatGPT, translate it into German and back into English — or generate in one language and translate to another — hoping the round trip scrambles the statistical fingerprint. Detection-evasion forums have recommended this for years. Does it work?

Honestly: obfuscation of this general type does degrade detection, and pretending otherwise would be dishonest marketing of the kind this site exists to avoid. In the same Weber-Wulff study, machine-paraphrased AI text — AI text transformed by another AI tool, which is what round-trip translation amounts to — dropped detection accuracy to about 26%. Across all obfuscation methods tested, roughly half of disguised AI text went undetected. Detectors are worst exactly where the stakes are highest.

But before anyone treats that as a recipe, three things are true at once.

First, it's a coin flip, not a cloak. A 26–50% detection rate means the method fails constantly. Anyone running this play across a semester of assignments is repeatedly spinning a chamber, and detectors are retrained on known evasion techniques. Turnitin has extended detection toward paraphrased and machine-modified text — paraphrase detection from July 2024, bypasser detection from August 2025, with no accuracy figure published for either. What slipped through in 2023 is training data now.

Second, the quality cost is real and visible. Round-trip translation is a lossy operation performed on text that was already statistically flat. It mangles idioms, breaks collocations, flattens register, and introduces the slightly-off phrasing that translators call translationese — "the possibility to" instead of "the chance to," articles dropped or duplicated, prepositions from the wrong language. Human readers notice this even when detectors don't. A professor who has read your in-class writing will notice fastest of all. The irony is sharp: the technique adds the very errors that make text look worse than honest AI output, in exchange for an evasion effect you can't rely on.

Third, it doesn't change what the work is. If a course bans AI-generated submissions, an AI-generated submission laundered through Estonian is still an AI-generated submission. The translation step is not a gray area; it's concealment, and concealment is usually what turns a policy violation into a formal misconduct finding. We build an AI humanizer ourselves, and our position on this is the same one printed across this site: it doesn't promise to beat any detector, because nobody can honestly promise that — and rewriting tools, ours included, are for editing your own permitted work into your own voice, not for disguising banned work as yours.

What we actually know about non-English detection

Most detector marketing is silent or vague about languages, so it's worth collecting what's on the record.

OpenAI said it plainly in the January 2023 announcement of its own classifier: the tool was "significantly worse" in non-English languages, and unreliable on short text. That classifier was retired seven months later with a 26% true-positive rate and a 9% false-positive rate — a story with its own lessons, which we tell in full in why OpenAI shut down its AI detector. Turnitin, for its part, states that its model was built and validated primarily on English long-form prose, with a roughly 300-word minimum. Several commercial vendors advertise multilingual support, but validated, per-language false-positive rates are almost never published; where a vendor claims specific non-English accuracy figures, treat them as unaudited self-reports until an independent test says otherwise. Turnitin, for one, states its own scope narrowly: English, Spanish and Japanese, with bypasser detection English-only.

The structural reason for the gap is unglamorous: training data. Detectors learn the statistical difference between human and machine text from large corpora, and those corpora are overwhelmingly English. A detector fed mostly English human writing has a thin model of what human Spanish, Czech, or Vietnamese prose looks like — and an even thinner model of what translated prose looks like, which is its own statistical dialect, distinct from both.

Here is the risk picture in one table:

ScenarioWhat the detector seesFalse-flag riskHonest assessment
Native English speaker writes in EnglishHuman statistical fingerprintLow (baseline)Detectors' best case; still not zero
Non-native speaker writes directly in EnglishHuman text with low perplexityRaised — Liang et al. found 61% average false flags on TOEFL essaysThe documented bias case
Human writes in another language, machine-translates to EnglishMachine-chosen English words carrying human ideasHigh — ~20-point accuracy drop in Weber-Wulff et al.The blind spot: honest work, machine fingerprint
Human self-translates, then edits heavilyMixed fingerprint, mostly humanModerateEditing in your own voice restores your statistical signature
AI text round-tripped through translation to evadeDegraded machine textDetection falls (~26% in one study) but is unpredictableUnreliable as evasion, corrosive to quality, and a policy violation where AI is banned
AI-assisted text in a course that permits AI, translatedMachine fingerprint (accurately)Not a "false" positive — but disclosure is what protects youCite the tools; the score stops mattering

Practical guidance for multilingual writers

None of this means multilingual writers should stop using translation tools. It means using them with eyes open.

Know your workflow's risk before someone else measures it. If your process involves machine translation and your text will face a detector — a class, a job application portal, an editor who screens submissions — assume the score will run high and plan accordingly. Running your own text through a checker first is not paranoia; it's the same instinct as proofreading. Our AI detector gives a sentence-level readout on up to 1,500 words per scan, free for 10,000 words a month, which at least shows you which sentences carry the machine fingerprint — though like every detector including ours, it reads statistical patterns, not authorship, and we publish no accuracy percentage for it because we haven't published a methodology that would make such a number meaningful.

Edit the translation until it's yours. The single most effective protection is also the best writing advice: treat machine translation output as a first draft. Rework sentence rhythms. Restore the idioms you actually use. Break up the uniform sentence lengths the engine produced. This isn't gaming the detector — it's reversing the exact thing the engine did to your voice, and the statistical fingerprint follows the voice.

Keep your process evidence. Drafts in your first language, translation-tool history, revision timestamps, version history in Google Docs or Word. If you are ever flagged, a documented trail from L1 draft to final English text is far stronger evidence than any counter-score. Accusations collapse against paper trails.

Disclose when disclosure is cheap. A one-line note — "drafted in Spanish and translated with DeepL, then revised" — costs nothing in most professional contexts and converts a potential gotcha into a described method. In academic contexts, ask the instructor before submitting, not after a flag.

Practical guidance for teachers of international students

If you teach multilingual students and use a detector, the evidence above should reshape how you read its output.

A high score on an international student's paper is consistent with at least three explanations: the student used AI in a prohibited way; the student wrote honestly in careful learner's English (the Liang pattern); or the student drafted in another language and translated (the Weber-Wulff pattern). The detector cannot distinguish these. Only process evidence can — drafts, notes, version history, and a conversation with the student about their text. Turnitin, to its credit, says the same thing: its own guidance frames the score as the start of a conversation, not a verdict, and its decision to display sub-20% scores as an asterisk rather than a number is a genuine acknowledgment that low-confidence output shouldn't be treated as evidence.

Two policy moves cost little and prevent most of the damage. First, decide — and tell students in writing — whether drafting in another language plus machine translation is acceptable in your course. Most instructors have never stated a position, and students are guessing. Second, never let a detector score be the sole basis for an accusation against any student, and be doubly cautious for multilingual students, where every published study shows the false-positive rate is highest. The Weber-Wulff authors, who ran the largest independent test, reached exactly this conclusion; so did Vanderbilt, which disabled Turnitin's AI indicator in August 2023 after doing the false-positive math at scale.

Where this leaves the translation question

Machine translation launders authorship in both directions, and detectors can't see through it either way. Human ideas come out wearing machine-chosen words and get flagged; machine ideas come out slightly scrambled and sometimes walk through. A tool that measures "how machine-typical is this English text" is answering the wrong question the moment a translation engine enters the workflow — and for millions of multilingual writers, a translation engine is always in the workflow.

The fix isn't a better score. It's remembering what scores are: statistical guesses about surface patterns, useful as a screening signal, useless as proof of anything about the person. Text isn't even the medium where detection works best — for a look at how differently the problem plays out with images, where cryptographic provenance is starting to do what statistics can't, see AI image detection vs. text detection.

FAQ

Will Google Translate or DeepL output get flagged as AI-generated? Often, yes. Neural translation engines generate every English word the same way a chatbot does, so the output carries a machine-typical statistical fingerprint. In the Weber-Wulff et al. (2023) study of fourteen detectors, accuracy on human-written text fell by roughly 20 percentage points once the text was machine-translated.

Is it against the rules to write in my language and translate my own work? Usually not — translating your own original ideas is not plagiarism and is not AI-generated content in the sense most policies mean. But some courses and journals have their own definitions, and a few explicitly restrict machine translation. Ask before you submit; a written answer from an instructor or editor is the cheapest insurance available.

Can translating AI-generated text back and forth really beat detectors? It reduces detection rates — one peer-reviewed study found machine-paraphrased AI text was caught only about 26% of the time — but it is unreliable, degrades the writing noticeably, and detectors are retrained against known evasion tricks. Where AI use is banned, translation laundering is concealment, which typically makes the misconduct case worse, not better.

Do AI detectors work in languages other than English? Much less well, as a rule. OpenAI said its own classifier was significantly worse outside English, and Turnitin states its model was built primarily on English prose. Vendors claiming multilingual accuracy rarely publish per-language false-positive rates, so treat those claims cautiously until independently tested.

I'm a non-native English speaker and I never use translation tools. Am I still at risk? The published evidence says yes. Liang et al. (2023) found seven detectors falsely flagged an average of 61% of human-written TOEFL essays, because careful learner's English is statistically predictable in the same way machine text is. Keeping drafts and version history is your best protection.

A detector flagged my translated document. What should I do first? Gather your process evidence before responding: the original-language draft, translation history, and revision timestamps. A documented trail from first-language draft to final text is stronger than any counter-score from another detector. Then explain your workflow plainly — translation of your own writing is a legitimate process you can describe without apology.

Should teachers ban machine translation to avoid this whole problem? That's a pedagogy question, not a detection question, but note what a ban costs: drafting in L1 and translating is a long-recommended strategy for multilingual writers, and banning it doesn't make detectors accurate on the learner's-English text students will produce instead. Clear disclosure rules solve more than bans do.

Key facts

  • Weber-Wulff et al. (International Journal for Educational Integrity, December 2023) tested 14 detection tools; accuracy on human-written text averaged 96%, but fell by roughly 20 percentage points on human text machine-translated from seven languages via DeepL and Google Translate.
  • The same study found machine-paraphrased AI text was detected only ~26% of the time, and about half of all obfuscated AI text went undetected; the authors concluded the tools are "neither accurate nor reliable."
  • Liang et al. (Patterns, 2023) found 61.22% of 91 human-written TOEFL essays were falsely flagged on average across seven detectors; 18 of 91 were flagged by all seven; the same detectors were near-perfect on native-speaker eighth-grade essays.
  • OpenAI's January 2023 classifier announcement warned the tool was significantly worse in non-English languages; OpenAI retired it in July 2023 at a 26% true-positive / 9% false-positive rate.
  • Turnitin states its detector was built and validated primarily on English long-form prose with a ~300-word minimum, and displays scores of 1–19% as an asterisk rather than a number.
  • Vanderbilt University disabled Turnitin's AI indicator in August 2023, citing false-positive risk at institutional scale.

Sources

  1. Weber-Wulff, D., et al. "Testing of detection tools for AI-generated text." International Journal for Educational Integrity 19, 26 (December 2023). link.springer.com/article/10.1007/s40979-023-00146-z
  2. Liang, W., et al. "GPT detectors are biased against non-native English writers." Patterns (Cell Press), 2023.
  3. OpenAI. "New AI classifier for indicating AI-written text." January 31, 2023, with July 20, 2023 retirement update. openai.com/index/new-ai-classifier-for-indicating-ai-written-text/
  4. Turnitin. AI writing detection FAQ / transparency documentation (English-primary training, ~300-word minimum, asterisk display for 1–19% scores).
  5. Vanderbilt University. "Guidance on AI detection and why we're disabling Turnitin's AI detector." August 2023.
All postsPublished by The HumanFlow team