humanflow

Why does ChatGPT use so many em dashes?

Nobody outside the model developers knows, and none of them has published an explanation — the observation that generated text uses em dashes heavily is widely reported and the cause is not documented anywhere we can point you to.

Last reviewed 15 August 2026 · The HumanFlow team

It is worth being blunt about the state of the evidence, because this question attracts a lot of confident answers. There is no study establishing the rate, no vendor statement explaining it, and no way for anyone outside a lab to inspect the training data or the preference tuning that would settle it. Anybody telling you exactly why is reasoning from the outside, and so is this page.

What can be said is what the punctuation mark does. An em dash is unusually flexible: it can stand in for a comma, a colon, a semicolon or a pair of brackets, and it is grammatical in more positions than any of them. A system choosing high-probability continuations is therefore choosing a mark that fits a great many slots, and one that carries no risk of being wrong the way a semicolon can be.

It is also common in the sort of text these systems are trained on. Edited books, magazine journalism and online essays use it far more freely than student writing or business email does, so prose modelled on published writing will contain more of it than the prose most people produce day to day.

That is an explanation of why the pattern is plausible, not evidence that it is the cause. Treat it as reasoning rather than as a finding, which is more than most pages on this question will tell you.

The practical part matters more than the explanation. No detector counts em dashes — they work from statistical patterns across a whole document rather than from a character frequency — so removing yours does not move a score. It does make your writing worse, because the em dash is correct punctuation that plenty of people use deliberately, and stripping it out to avoid suspicion means letting a rumour edit your prose.

When this answer changes

Model behaviour shifts between versions, and several developers have publicly adjusted style in response to exactly this kind of complaint. An observation about output in one year may not describe the next.

The pattern is much weaker when the prompt specifies a house style or supplies examples to follow, because both push the output away from the default register.

Where to go next

Sources

One of our direct answers.