
Amazon is buying rare books, destroying them by cutting their spines, and scanning them to train its AI models at a Las Vegas facility, 404 Media reported by tracking a rare book to the site.
The company values rare and out-of-print texts because they were written before 2022 and thus contain no AI-generated content—a key advantage, since LLMs that ingest too much AI-generated text risk suffering 'model collapse,' a decline in output quality.
What happened
Amazon is purchasing rare books through commercial channels, cutting off their spines, and scanning them at its facility in Las Vegas (identified as VGT3, marked with a dinosaur logo) to use as training data for its AI models, according to 404 Media, which tracked a rare book to confirm the practice.
Why it matters
Rare books—especially out-of-print texts unavailable online—are valuable training sources because they predate 2022 and thus were not written by AI models. Companies like Amazon face a technical problem: when LLMs train on AI-generated text, they risk 'model collapse,' a degradation in output quality, so human-written historical texts offer cleaner training material than data already contaminated by earlier AI outputs.
What to watch
Amazon stated it 'purchases books through commercial channels to improve the products and services customers use,' but the practice raises questions about which rare collections are being targeted and whether the destruction of scarce cultural artifacts for commercial AI training will face regulatory or legal scrutiny.
Ask the AI about this article →
Amazon's shift to rare books reflects a scaling challenge now facing all major AI companies: the internet has been largely exhausted as a training source. As the article notes, LLMs have "already ingested what they can from the internet," leaving companies like Amazon searching for fresh text at scale. Rare and out-of-print books represent a new frontier because they are both scarce (and thus not widely known to be in circulation) and historically pure—guaranteed to predate the AI era.
The technical rationale is sound: LLMs trained on AI-generated outputs suffer from what researchers call "model collapse," a degradation in reasoning and output quality. This creates a direct incentive to pursue human-authored texts, especially older ones where authorship is unambiguous. The article also references Anthropic's admission of training on "illegally pirated books," suggesting that outright unlicensed acquisition is already a competitive pressure in the space. Amazon's purchasing through "commercial channels" may represent a more legally defensible approach, but the wholesale destruction of rare texts raises a distinct question: whether the preservation value of these artifacts—to libraries, scholars, and cultural institutions—outweighs their utility as training data.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.