AIToday

AI can now link pseudonymous posts by writing style alone

LessWrong AI13h ago
AI can now link pseudonymous posts by writing style alone

Key takeaway

An essayist argues that language models can now identify authors across pseudonymous identities by analyzing the statistical fingerprint of their writing style. While a public tool to match new text against all a person's prior pseudonymous work does not yet exist, the author believes it will soon become feasible as LLMs improve, posing a threat to anonymous and pseudonymous online expression.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    An essayist describes a conjecture that large language models can identify authors across pseudonymous identities by analyzing statistical patterns in their writing. The author notes that LLMs have grown capable enough that Claude 4.8 successfully identified the author from the first 1,000 words of a draft post.

  • Why it matters

    If text-based author identification becomes widely available as a public tool, it would undermine pseudonymous blogging and online anonymity. The author suggests such a website—one that matches new text to all prior writing by the same person across different names—does not yet exist on the public internet but suspects it will become possible and easy to build soon.

  • What to watch

    The author wrote most of the essay in mid-2025 and notes that LLM capability in author identification has already improved significantly since then, suggesting the technical barrier to such a tool is dropping.

In Depth

The author opens with a straightforward conjecture: any writer who publishes a significant body of text under different names leaves a statistical fingerprint in the prose itself—patterns invisible to human readers but extractable by machine learning. The thought experiment is vivid: imagine a website where you paste new text and receive back links to everything that writer has published under any pseudonym. The author acknowledges the tool would be imperfect but believes it could work reliably enough to matter.

The author explicitly states that no such website exists on the public internet as far as they know. However, they suspect it is possible to build and that it will become easy to do so soon. The implication is clear: pseudonymous blogging, which has long relied on the assumption that different names represent different people unless proven otherwise, would lose that protection.

Critically, the author notes they drafted the essay in mid-2025, then shelved it for a year while refining technical material. During that interval, they observe that large language models have improved significantly at author identification. They provide a concrete example: Claude 4.8, given the first 1,000 words of a draft of their own post, correctly identified them as the author. This real-world validation of the core claim—that an LLM can identify an author from partial text—underscores that the theoretical conjecture is already becoming practical reality.

Context & Analysis

The essay is rooted in a observation about the inherent statistical properties of individual writing: every person leaves a distinctive fingerprint in their prose—word choice, sentence rhythm, argument structure, and topic patterns that accumulate across texts. The author's core claim is that this fingerprint is extractable by machine learning, particularly by large language models, making it possible to link disparate pseudonymous accounts to a single author.

The timing of the piece is significant. The author drafted it in mid-2025 but delayed publication while refining theorem statements. In that intervening year, LLM capabilities in authorship attribution have improved materially—the fact that Claude 4.8 can identify the author from a partial draft suggests the technical hurdle is already lower than when the conjecture was first formed. This narrowing gap between the theoretical possibility and practical capability is the real tension the author highlights: the tool does not yet exist as a public service, but the underlying capability is demonstrable and the barrier to deployment appears to be shrinking.

FAQ

Does a public tool to de-anonymize writers already exist?
No. The author states that as far as they know, no such website exists on the public internet, though they suspect it will become easy to build soon.
How well can current AI identify authors from their writing?
The author reports that Claude 4.8 successfully identified them as the author from the first 1,000 words of a draft of their essay.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime