AIToday

Why AI's Em Dash Habit Reveals Unedited AI Output

Hacker News1d ago

Key takeaway

Large language models have developed a telltale habit of overusing em dashes in their output—a punctuation mark that humans rarely type because it requires a special keyboard shortcut. Models learned this pattern from their training diet of professionally edited text, where dashes cost nothing to produce. The overuse has become a visible fingerprint of AI-generated content pasted unchanged into workplace communication, signaling to readers that the sender outsourced judgment rather than engaging with the message themselves. The article warns that this pattern, multiplied across a team, erases individual voice and workplace culture.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Large language models use em dashes—a punctuation mark traditionally rare in casual writing—far more frequently than human writers do, because they trained on professionally edited text (books, academic papers, journalism) where the mark appears constantly and costs nothing to type with a keyboard shortcut.

  • Why it matters

    The overuse of em dashes has become a recognizable fingerprint of AI-generated text pasted unchanged into Slack messages, Notion pages, and other workplace communication. Unlike humans, who notice their own writing habits and edit them out, models generate one token at a time with no memory of previous uses and no training to recognize or break their own patterns. When readers spot dash-heavy text, they perceive the sender as having outsourced judgment rather than standing behind the message—a small signal of care that affects trust and the sender's reputation.

  • What to watch

    The article urges readers to read AI output before sending it and rewrite parts that don't sound like them. Wholesale pasting of model output without editing contributes to team homogenization—flattening individual voice and workplace culture into a single polished, hedged cadence that erases the 'shortcuts, register, and little quirks' that give a team its character.

In Depth

The em dash is a character almost nobody used to type. Named for its width—traditionally one em, the width of the letter M—it has a much shorter cousin, the en dash. Printers adopted something close to it in English drama as far back as 1588, in Maurice Kyffin's translation of Terence's Andria, using a run of hyphens to mark broken, interrupted speech. A few decades later, other translations of the same play used the same trick dozens of times. Eventually the expensive folio editions settled on a single unified dash instead of a chain of hyphens, mostly because it looked better. From there the em dash moved comfortably into professionally edited prose—novels, essays, journalism—where it appeared in the Chicago Manual of Style and the AP Stylebook. American convention kept it unspaced; British and Canadian style guides often preferred a spaced en dash instead. But in everyday, unedited human writing, the em dash stayed rare because it was a small pain to produce. A standard keyboard gives you a hyphen and nothing else. Getting a real em dash meant Option+Shift+Hyphen on a Mac, an alt code on Windows, or the old Courier workaround of stringing two hyphens together, because monospaced typewriter fonts never had a dedicated glyph for it.

Large language models learned to write from the opposite end of that spectrum. Their training diet came from books, academic papers, and edited journalism—exactly the domains where the em dash shows up constantly, because it costs nothing to produce with a keyboard shortcut and carries none of the friction that keeps casual writers away from it. That training taught the models that this is what careful, "considered" writing looks like, so they reach for the same device sentence after sentence, whether or not the content actually calls for a pause or an aside. What the models don't have is what a human editor has: a sense of their own tics. A writer eventually notices they lean on a phrase too much and cuts it. A model generates one token at a time, with no memory of how many times it already used the same trick three paragraphs up. Nobody trained it to notice its own habits, so nothing ever tells it to stop.

The result is that the em dash has stopped being neutral punctuation. A person typing quickly in a DM doesn't go out of their way to produce a character that needs a special key combination. They reach for a hyphen, a comma, or nothing at all. A message full of unspaced dashes doing the exact same rhetorical move over and over isn't how people type under normal conditions. It looks exactly like what it is: text a model produced, pasted in unchanged. This creates a problem rooted in trust. A Slack DM or Notion page written entirely in that voice reads like nobody actually stood behind it. If the sender didn't write it themselves, did they even think it through, or did they just ask a model what to say and forward the reply? The message stops being proof of someone's judgment and starts being proof that the judgment got outsourced. Writing something yourself, even a couple of sentences, is a small act of care toward whoever's reading it. Pasting model output wholesale, dashes and all, without even a pass to make it sound like you, tells the reader you didn't think the message was worth that much of your time. And underneath both of those, slower and easier to miss day to day, sits homogenization. When enough people on a team lean on the same tool for the same tone, every DM, every project update, every Notion doc starts sounding like it came from the same person, and that person is nobody in particular. A team's culture lives partly in how people actually talk to each other: the shortcuts, the register, the little quirks of individual voice. Flatten all of that into one polished, hedged, dash-heavy cadence, and you lose something that was quietly doing real work. The recommendation is to read what comes back before it goes anywhere, and rewrite the parts that don't sound like you. A tool that ghostwrites your Slack messages and your Notion pages is, quietly, a tool writing your reputation too.

Context & Analysis

The em dash occupies a peculiar place in modern writing history. For centuries it belonged to professional typesetters and editors—people whose job was to polish prose. It first appeared in English drama around 1588 in Maurice Kyffin's translation of Terence's Andria, where it marked broken, interrupted speech. From there it settled comfortably into novels, essays, journalism, and style guides like the Chicago Manual of Style and AP Stylebook. The barrier to everyday writers was simple: a standard keyboard offered no em dash. On a Mac you needed Option+Shift+Hyphen; on Windows, an alt code; or you could use the old typewriter trick of two hyphens in a row. That friction kept the mark rare in unedited, casual human writing.

Large language models encountered the em dash from the opposite end of that spectrum. Their training data—books, academic papers, edited journalism—is exactly the domain where dashes are abundant and free to produce. The models learned that this punctuation pattern is what "careful, considered" writing looks like, so they reach for it repeatedly. But they lack what a human editor has: a sense of their own tics. A writer notices they lean on a phrase too much and cuts it. A model generates one token at a time, with no memory of how many times it used the same trick three paragraphs earlier, and no training to recognize or break the pattern. The result is text that reads like nothing a human would produce under normal conditions—a visible fingerprint of AI authorship.

FAQ

Why do AI models use em dashes so much more than people do?
Models trained on professionally edited text—books, academic papers, journalism—where em dashes appear constantly because they cost nothing to produce with a keyboard shortcut. Humans rarely use them in casual writing because a standard keyboard requires a special key combination (Option+Shift+Hyphen on Mac, an alt code on Windows, or a workaround with two hyphens).
How can you tell if text is AI-generated based on em dashes?
A person typing quickly in a DM or message wouldn't go out of their way to produce an em dash multiple times; they'd reach for a hyphen, comma, or nothing. Text with unspaced dashes doing the same rhetorical move over and over—the exact same pattern throughout—looks like what it is: output from a model that generates one token at a time and has no memory of how many times it already used the same trick, so nothing tells it to stop.
What does the article recommend doing with AI-generated text?
Read what comes back before sending it, and rewrite the parts that don't sound like you. Pasting model output unchanged, dashes and all, without a pass to make it sound like your own voice, tells the reader you didn't think the message was worth your time.

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No discussion yet for this article

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime