humanflow
Writing · The HumanFlow team · 15 min read

Writing in your own voice when AI drafts it

AI drafts fast and flattens voice. Here is what voice is made of mechanically, why models erase it, and the five edits that put it back — with the before and after.

AI is a fast first-drafter and a reliable flattener. It smooths your quirks into the same even, tidy, slightly hollow prose everyone else's model produces. The result reads fine and sounds like no one.

Keeping your voice is not a matter of tricks or of a better prompt. It is a matter of putting a few of your own decisions back into text that had them removed. This post is about which decisions those are, why a model removes them in the first place, and what the edits look like on the page.

One thing this is not about: getting a draft past a detector. We do not publish bypass advice and we do not claim our own tool defeats detection — the reasoning is in our methodology. What follows is craft advice. That some of it also happens to be what detectors read as human is a consequence, not the objective, and treating it as the objective is how people end up writing worse.

What a model is actually doing to your voice

It helps to know the mechanism, because it explains why the same three problems come back no matter how you word the prompt.

A language model generates text by repeatedly choosing a likely next token. "Likely" is the operative word. Averaged over a colossal amount of text, the likely next word after "In conclusion," is not the word you would have chosen — it is the word most writers chose. The model is not trying to sound like anyone in particular. It is trying to sound like the centre of a distribution, and the centre of a distribution has no accent.

That is the whole problem in one sentence, and it produces three specific symptoms.

Predictability. The technical name is perplexity — roughly, how surprised a language model is by each next word. Text that goes exactly where a model expects has low perplexity. Your writing has higher perplexity than a model's default output not because you are unpredictable in some romantic sense, but because you know things the average of the internet does not: which example you saw last week, which word your field actually uses, which objection your reader will raise.

Even rhythm. Human writing runs in bursts — a short jab, then a long winding sentence that doubles back on itself, then a short one again. Model output tends to hold a medium stride and keep it for paragraphs. The measure of that variation is burstiness, and a flat rhythm is the single most audible tell in machine prose. You can hear it before you can name it.

Hedging by default. Models are trained to be helpful and inoffensive, which in practice means trained to qualify. May, might, can be, is often considered, it is important to note. Each individual hedge is defensible. Forty of them in a thousand words is a writer with no position, and hedging at that density reads as evasion whether or not it was meant that way.

None of these is a flaw you can prompt away entirely, because none of them is a bug. They are what "sound like everyone" looks like from the inside.

Start from the draft, not the blank page

Let the model get you to a rough draft, then treat it as raw clay rather than a finished piece. Move fast on structure and slow down on voice.

This ordering matters more than it sounds. The common failure is to accept a draft's structure because its sentences read smoothly, and smooth sentences are exactly what a model is best at. A draft can be fluent and still argue nothing, sequence badly, or bury the point in paragraph four. Fix that first, while you are still willing to delete things.

A practical test: read only the first sentence of each paragraph, in order. If that sequence does not tell the story, no amount of sentence-level editing will save the piece, and you are about to spend an hour polishing prose you should be cutting. This is also the argument in rewriting versus editing it yourself — a rewriting tool is a late-stage instrument, and a structural problem is not what it is for.

Cut the filler

AI loves connective throat-clearing: "In today's fast-paced world…", "It is important to note that…", "Moreover", "Furthermore", "That being said". Delete it. Your reader never needed the runway.

This is the highest-yield edit available and the least interesting to make, which is why it gets skipped. It is also the most mechanical, so it is the one worth using a tool on. Our AI word finder tracks 109 terms across five families — connective throat-clearing, abstract stand-ins, hedges and softeners, inflated verbs and adjectives, and stock openers and closers — and highlights them in a draft so you can see the density rather than guess at it.

Two cautions, because a word list is easy to misuse.

These words are not forbidden. However is a fine word. Moreover has a job. The signal is density and predictability, not presence — a paper that uses furthermore twice is a paper; a paper that opens every third paragraph with it is a template. Deleting every instance produces prose that is choppy in a new and equally artificial way.

Cutting the word is not the edit. Cutting the word is the start of the edit. "It is important to note that the sample size was small" becomes "The sample was small" — you removed seven words and, more to the point, you stopped telling the reader what to find important and let the fact do it.

The other half of the same problem is nominalization: verbs turned into nouns. "We conducted an analysis of" instead of "we analysed". "The implementation of the policy resulted in an improvement" instead of "the policy improved". Nominalized prose sounds official and dead at the same time, and it is everywhere in model output because it is everywhere in the institutional text models were trained on. The nominalization finder catches these, and the fix is nearly always to find the buried verb and let it do the work.

Let sentences vary

Real writing has rhythm. Model output tends toward one length. Break it up.

The best diagnostic costs nothing: read it aloud. Not silently — aloud, or at least subvocalised. Your ear catches the metronome long before your eye does, and it catches it in your own writing, where you are otherwise blind. If it sounds like a form letter, it is one.

Then look at the shape. The sentence rhythm checker plots sentence lengths across a passage so you can see the flatness rather than intuit it, and the sentence opener checker catches the related tic — six paragraphs in a row starting with The, or with a participial phrase, or with the subject in exactly the same position.

What varying rhythm actually means in practice:

  • Put a short sentence after a long one. Not everywhere. After the sentence that did the heavy lifting, so the point lands.
  • Let one sentence be genuinely long and earn it — a sentence that accumulates, that adds a clause because the thought genuinely has another part to it, rather than one padded to look substantial.
  • Start a sentence with a conjunction if that is how the thought connects. But is a sentence opener. So is And. The rule against them was invented for children and never applied to good prose.
  • Use a fragment when the emphasis is worth the grammar. Sparingly. Like that.

The point is variation, not brevity. Uniformly short sentences are as flat as uniformly medium ones, and the advice to "write short sentences" has produced as much bad prose as the advice to write long ones.

Keep one specific thing

A number, a name, a date, a small true detail. Specifics are the fingerprints a model smooths away, and they are what make writing sound lived-in.

This is the edit that most improves a draft and the one no tool can do for you, because the specifics are not in the text — they are in your head. A model cannot tell you which student asked the question, what the meeting actually cost, or which of two similar tools you tried first and abandoned. It will happily write "many educators have expressed concern", which is true of everything and evidence of nothing.

Replace an abstraction with a particular whenever you can:

Draft: Recent studies have shown that AI detection tools can produce false positives, which has raised significant concerns among educators regarding their use in academic settings.

Edited: Weber-Wulff and colleagues tested fourteen detectors and found the worst flagged half the human-written samples in their test set. Half.

The second is shorter, checkable, and has a human deciding what mattered — including the decision to repeat the word for emphasis, which is a choice a model averaging a distribution will not make. It also carries a source, which is what turns a claim into something a reader can act on. (That figure is real; the study is in our literature review, and the reason we keep repeating it is on what detector accuracy means.)

If you genuinely have no specific to add, that is worth knowing too. It usually means you are writing about a topic you have not thought about yet, and no editing pass fixes that.

One paragraph, worked through

Abstract advice is easy to agree with and hard to apply, so here is a whole paragraph of the kind a model produces on request, and what happens when the five edits go through it.

Draft. In today's rapidly evolving educational landscape, it is important to note that AI detection tools have become increasingly prevalent. Moreover, these tools utilise sophisticated algorithms in order to make a determination regarding whether content was generated by artificial intelligence. However, it is worth noting that there are significant concerns regarding their accuracy. Consequently, many educators have expressed reservations about their implementation, and institutions may wish to carefully consider a range of factors before adopting such systems.

Eighty-one words, five sentences, and it says almost nothing. Every sentence is between fourteen and twenty-two words. Four of the five open with a connective — In today's, Moreover, However, Consequently. There is one verb doing real work in the entire paragraph (utilise, and even that is standing in for use).

Edited. Universities bought AI detectors faster than anyone tested them. The tools work by measuring how statistically predictable your prose is — not by finding evidence of anything. That distinction matters, because predictable is not the same as machine-written, and the largest peer-reviewed test of fourteen detectors found the worst of them flagged half the human samples it was given. Half. Faculty noticed. Several universities have since turned the feature off.

Seventy-two words, six sentences, and now something is being claimed. Look at what changed:

  • Sentence lengths run 9, 20, 38, 1, 2, 8. The draft ran 14–22 throughout. The long sentence earns its length by carrying the actual argument; the one-word sentence after it is the whole point of having varied rhythm at all.
  • The connectives are gone, and nothing replaced them. Moreover and Consequently were doing no logical work — the sentences already followed from each other, which is what made the connectives skippable.
  • The nominalizations are unpicked. "Make a determination regarding" became work by measuring. "Have expressed reservations about their implementation" became noticed.
  • Two specifics arrived: fourteen detectors, half the human samples. Both are real and both are sourced on this site.
  • It takes a position. "Bought faster than anyone tested them" is a claim someone could dispute. The draft's "may wish to carefully consider a range of factors" is not.

The edited version is shorter. That is typical, and it is the thing people find most surprising: restoring voice is mostly subtraction, and the additions are specifics rather than words.

Say the thing you actually think

The hardest of the five, and the one that separates prose with a voice from prose that has merely been de-robotified.

Models hedge because hedging is safe. The cumulative effect is writing that surveys a topic without ever committing to a position on it — every consideration acknowledged, no consideration weighted. Readers can feel this even when they cannot name it, and it is why so much competent AI-assisted writing is unmemorable rather than bad.

The edit is to find the sentence where you almost said something and let it say it:

Draft: There are arguments on both sides, and institutions may wish to consider a range of factors when developing their approach.

Edited: An institution that acts on a detector score without looking at the student's drafting history is going to be wrong about somebody, and it will be wrong about the same kinds of student every time.

That is a position. It can be argued with, which is precisely what makes it worth reading. Notice that it is not less careful than the hedged version — it is more careful, because it says something specific enough to be checked. Vagueness is not caution. It is the appearance of caution with none of the work.

How this goes wrong

The advice above has a failure mode, and it is worth naming because it is common and it produces writing that is worse than the flat draft it replaced.

Over-correction. Told that AI writing has even rhythm, people write in a deliberately jagged one. Told that it hedges, they strip every qualifier and end up asserting things they do not know. Told to add specifics, they add decoration — a number that is not load-bearing, a name that illustrates nothing. The result reads as someone performing personality rather than having one, which is a more irritating failure than blandness because the reader can see the effort.

The corrective is that every one of the five edits is in service of the reader, not of sounding human. Vary rhythm so emphasis lands where you want it. Cut hedges you do not mean, and keep the ones you do — a genuine uncertainty stated plainly is better writing than false confidence. Add specifics that carry the argument. If an edit does not make the piece clearer or more useful, it is costume.

Editing for a score. The other failure is treating a detector's output as the target. It is a poor target: detectors disagree with each other on identical text, they change without notice, and a number that goes down tells you nothing about whether the prose got better. People who edit toward a score tend to produce text that is strange in ways they cannot explain, because they are optimising against a measurement they cannot see inside. Edit for a reader. Check a score afterwards if you want, and treat a disagreement between two detectors as information about detectors.

A caveat that belongs here

Most of this advice assumes something it should state: that you have enough command of English to break its conventions on purpose.

That assumption does not hold for everyone, and it is not a small point. Liang and colleagues found that detectors misclassified non-native English writers' work as machine-generated at dramatically higher rates than native writers' — the reason being that writing which is careful, conventional and vocabulary-constrained looks statistically like model output, whether a person or a model produced it. The same properties that make a second-language writer's prose competent make it read as predictable.

So "write with more voice" is advice that costs different people different amounts. A native speaker breaks a rule and it reads as style. A second-language writer breaking the same rule may simply be marked down for it, by a human or by a classifier. Anyone giving this advice — including us — should be honest that it is easier to follow in your first language, and anyone acting on a detector score should know that the population most likely to be flagged is the one least able to write their way out of it. We go through the evidence on why non-native writers are flagged more often.

The craft advice still stands. Specifics, rhythm and a stated position improve writing in any language. But the framing "sound less like a machine" quietly assumes the problem is that you are writing like a machine, when for a large number of people the problem is that a machine is reading them badly.

Where the tools stop

Worth saying plainly, since this is our blog and we sell one of these.

A rewriting tool is a late-stage instrument. It does a mechanical pass on prose that is already correct — it will vary your rhythm, drop the throat-clearing, and unpick some nominalizations. It will not find the specific example you did not include, it will not take a position you did not take, and it will not fix a draft whose structure is wrong. We have written up where a rewrite stops and how it compares with editing the thing yourself, and the honest summary is that doing the five things above by hand will get you further than any tool will, including ours.

If you want to see what a detector reacts to before you touch anything, our detector reports a score with the signals behind it and how many sentences fell in each band. Read it as a rhythm-and-predictability report rather than a verdict, because that is what it measures. The flat, evenly-weighted lines it highlights are usually the ones worth rewriting on their own merits — which is the only reason to rewrite them.

And the disclosure point, which no editing changes: if your institution or your client requires you to say that AI was involved, say so. A draft that reads like you does not become a draft you wrote unassisted, and the obligation is about how the work was produced, not how it sounds. We keep a disclosure generator for exactly that, and our position on the whole question is at is using an AI humanizer cheating.

The five edits, in order

  1. Fix the structure first. Read the first sentence of every paragraph in sequence. If that is not the argument, stop editing sentences.
  2. Cut the filler, then rewrite what is left. Removing "it is important to note that" is step one; the sentence underneath usually wants rebuilding around its verb.
  3. Vary the rhythm. Read it aloud. Put a short sentence where the long one landed.
  4. Add one specific per section. A number, a name, a date, a thing that happened.
  5. Say what you think. Find the hedged sentence and let it commit.

None of this is about defeating a classifier. It is what editing has always been, applied to a draft that arrived with the editing already undone. The reason it matters more now is only that the first draft is cheaper than it used to be — which means the difference between writing that is worth reading and writing that is not has moved almost entirely into the second one.

All postsPublished by The HumanFlow team