
An investigation by 404 Media, using an AirTag planted in a rare book, has confirmed that Amazon is purchasing bulk lots of rare books and destroying them at a Las Vegas facility to train its AI models.
The facility's team scans and discards works that booksellers value for their historical or intellectual significance, though Amazon does not target the highest-priced first editions but rather older, lower-value books—many never translated or widely distributed.
Amazon declined to comment on the findings beyond stating it purchases books through commercial channels to improve its products.
What happened
An AirTag hidden in a rare book tracked its path to an Amazon facility in Las Vegas where a team systematically tears books from spines, scans pages, and destroys them for AI training data. The warehouse door featured a logo of a Tyrannosaurus rex preparing to devour a book.
Why it matters
Amazon is sourcing rare books—works with historical, intellectual, or sentimental value that booksellers carefully assess—and destroying them after scanning to gain training data advantage over rivals like Google, OpenAI, and Anthropic. Both competitors have publicly stated they do not train on rare or antique books, whereas Amazon appears to see such texts as a source of unique, unguarded content.
What to watch
The facility's workers reported in forums that Amazon ran low on books earlier this year and worried the warehouse might shut down without fresh supplies. Amazon now appears to be systematically targeting books by ISBN to ensure maximum unique works in its training datasets—suggesting the destruction of rare books may accelerate.
Ask the AI about this article →
For over a year, booksellers have suspected that AI firms were bulk-buying rare books and destroying them for training data, but confirmation has been elusive until 404 Media's investigation. The evidence that Amazon is behind at least some of these orders underscores a hard trade-off in the race for frontier AI models: the need for massive volumes of unique, original text versus the destruction of cultural artifacts. Amazon's statement that it "purchases books through commercial channels to help develop and improve the products and services our customers use" does not specify AI training, yet the facility's operational reality—workers scanning and destroying books at an industrial pace—speaks to the scale of data consumption required to compete with leading firms like Google, OpenAI, and Anthropic.
The targeting strategy appears methodical. Workers were trained to scan ISBNs before processing books, which booksellers theorize is an effort to work systematically through a comprehensive list of printed books. When the facility experienced a shortage of supply earlier this year, workers worried the warehouse might shut down entirely, suggesting that the volume of rare books flowing through is essential to Amazon's roadmap. This pressure to feed the AI training pipeline may be why Amazon is not filtering for the highest-value first editions—those would attract legal or reputational scrutiny—but rather older, obscure works that, while possessing historical or intellectual value to specialists, carry low commercial price tags and pass through the marketplace with minimal fanfare.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.