Humanizing AI text as a non-native English writer
Vary your sentence length before anything else. AI detectors flag second-language writing at high rates because both share a signature: common vocabulary and even, predictable structure. In peer-reviewed testing, seven detectors misclassified 61.22% of TOEFL essays as AI-generated. Varying rhythm addresses the actual signal.
Typical length · Any length. This is about pattern, not word count. · Last reviewed 16 August 2026
Before and after
Before
The research shows that exercise is good for health. Many studies have found this result. It is important for people to exercise regularly. Doctors recommend thirty minutes each day.
After
Exercise is good for you — that much is settled. What surprised me in the research was the threshold: thirty minutes a day, and most of the benefit is already there. You do not need the gym membership you feel guilty about.
Four short declaratives of similar length became one short, one long and one medium. A personal reaction and a specific number replaced generic reporting. Neither version has grammar errors — the first simply reads as machine-like.
Why detectors flag second-language writing so often
This is the best-measured failure in the whole field. In peer-reviewed testing published in Patterns, seven detectors classified 61.22% of TOEFL essays written by non-native English speakers as AI-generated — 89 of the 91 essays tripped at least one tool — while essays by native-speaking US eighth-graders were handled almost perfectly.
The mechanism is not prejudice in the software. Detectors measure how predictable text is, and writing in a second language under exam conditions is predictable for entirely honourable reasons: you reach for constructions you are confident in, you choose the common word over the unusual one, and you keep sentences to lengths you can control. That is what competence under pressure looks like, and it produces the same statistical signature as machine output.
So the flag is not evidence that the writing is bad. Often it is evidence the opposite — that the writer was careful.
Vary length before you vary vocabulary
The instinct after being flagged is to reach for more elaborate words. It is the wrong move twice over: unusual vocabulary chosen from a thesaurus reads as strained to a human marker, and it leaves the sentence rhythm — the thing detectors actually respond to — completely unchanged.
Sentence length is both the stronger signal and the easier thing to control deliberately. Look at the before and after above: the vocabulary barely moves. What changes is that four sentences of roughly equal length become one short, one long and one medium, and the paragraph acquires a shape.
A practical version, if you want a rule to apply: after writing a paragraph, read only the sentence lengths. If they are all within a few words of each other, join two and cut a third. You do not need better English to do that — you need to be willing to write one sentence that runs long and one that stops early.
Add the thing a model could not know
The second change in the example is a specific number and a personal reaction: the thirty-minute threshold, and the writer being surprised by it. Neither is decoration. A generic model has no reason to generate either, because it has no relationship to the material.
This is the part that survives every change in detection technology, because it is not a trick. Writing that contains something only you could have contributed — a figure you looked up, a case from your own reading, an objection you had — is harder to produce mechanically and better on its own terms.
It is also the strongest thing to have if the work is ever questioned. Specific, checkable content is evidence about how the document was made in a way that fluent generic prose is not.
If you have already been flagged
Do not rewrite the submitted work to make it score better — that changes the document under dispute and is very hard to explain afterwards. The evidence that helps is about how the work was produced, not how it reads.
The measured bias is a legitimate part of that argument, and it is worth citing precisely. The 61.22% finding is a seven-detector average from a study that did not include Turnitin, so quoting it against a Turnitin result invites a correction. Cite it for the general point, and say that is what you are doing.
Formats with the same problem
Essays — the format most often flagged, and how to vary rhythm without losing accuracy.
Personal statements — where a second-language writer is judged on voice and has the most to lose.
Discussion posts — the format where one specific detail carries the whole submission.
Questions
- Do AI detectors really flag non-native English writers more often?
- Yes, and it is measured rather than anecdotal. Seven detectors flagged 61.22% of TOEFL essays by non-native English speakers as AI-generated in peer-reviewed testing published in Patterns, while classifying native-speaker essays almost perfectly. Turnitin's own funded study reports near-identical rates for both groups above 300 words, so the picture is contested — both findings are on our accuracy page.
- Should I use more advanced vocabulary to avoid being flagged?
- No. Thesaurus substitution leaves sentence rhythm untouched, which is the signal detectors respond to most, and it reads as strained to a human marker. Varying sentence length does more for both audiences and requires no vocabulary you do not already have.
- Is it unfair to ask me to change how I write?
- Yes, and that is worth saying plainly rather than working around. The burden falls on second-language writers because of how these tools work, not because of anything those writers did. We publish the research on that bias even though it undercuts the category we sell into.
- Will running my work through a humanizer fix this?
- It may lower a score and it is not a solution to the underlying problem, and we will not claim otherwise. If your institution requires disclosure of AI assistance, that obligation is unchanged. If you have been wrongly accused, drafts and version history are worth more than any rewrite.
What to watch for
- If your work has been wrongly flagged, keep your draft history. It is the strongest evidence available.
- Do not let a rewrite introduce idioms you would not use — that reads as inauthentic to a human marker even if it satisfies a detector.
If your writing gets flagged
Rewriting for rhythm and specificity tends to lower detection scores, because that is what detectors read as human. It is not a guarantee — detectors disagree with each other and change without notice, and we do not promise a result from any of them.