AIToday
Ars Technica AIPublished: Aug 18, 2026, 04:00 JST3 min read

Amazon destroying rare books to train AI, tracker reveals

Amazon destroying rare books to train AI, tracker reveals

Key takeaway

  • An investigation by 404 Media, using an AirTag planted in a rare book, has confirmed that Amazon is purchasing bulk lots of rare books and destroying them at a Las Vegas facility to train its AI models.

  • The facility's team scans and discards works that booksellers value for their historical or intellectual significance, though Amazon does not target the highest-priced first editions but rather older, lower-value books—many never translated or widely distributed.

  • Amazon declined to comment on the findings beyond stating it purchases books through commercial channels to improve its products.

3 Key Points

  1. What happened

    An AirTag hidden in a rare book tracked its path to an Amazon facility in Las Vegas where a team systematically tears books from spines, scans pages, and destroys them for AI training data. The warehouse door featured a logo of a Tyrannosaurus rex preparing to devour a book.

  2. Why it matters

    Amazon is sourcing rare books—works with historical, intellectual, or sentimental value that booksellers carefully assess—and destroying them after scanning to gain training data advantage over rivals like Google, OpenAI, and Anthropic. Both competitors have publicly stated they do not train on rare or antique books, whereas Amazon appears to see such texts as a source of unique, unguarded content.

  3. What to watch

    The facility's workers reported in forums that Amazon ran low on books earlier this year and worried the warehouse might shut down without fresh supplies. Amazon now appears to be systematically targeting books by ISBN to ensure maximum unique works in its training datasets—suggesting the destruction of rare books may accelerate.

Ask the AI about this article →

Context & Analysis

For over a year, booksellers have suspected that AI firms were bulk-buying rare books and destroying them for training data, but confirmation has been elusive until 404 Media's investigation. The evidence that Amazon is behind at least some of these orders underscores a hard trade-off in the race for frontier AI models: the need for massive volumes of unique, original text versus the destruction of cultural artifacts. Amazon's statement that it "purchases books through commercial channels to help develop and improve the products and services our customers use" does not specify AI training, yet the facility's operational reality—workers scanning and destroying books at an industrial pace—speaks to the scale of data consumption required to compete with leading firms like Google, OpenAI, and Anthropic.

The targeting strategy appears methodical. Workers were trained to scan ISBNs before processing books, which booksellers theorize is an effort to work systematically through a comprehensive list of printed books. When the facility experienced a shortage of supply earlier this year, workers worried the warehouse might shut down entirely, suggesting that the volume of rare books flowing through is essential to Amazon's roadmap. This pressure to feed the AI training pipeline may be why Amazon is not filtering for the highest-value first editions—those would attract legal or reputational scrutiny—but rather older, obscure works that, while possessing historical or intellectual value to specialists, carry low commercial price tags and pass through the marketplace with minimal fanfare.

FAQ

How was Amazon's book destruction confirmed?
A bookseller planted an AirTag in a rare book that was part of a bulk order. The tracker led 404 Media to an Amazon facility in Las Vegas where a team was documented tearing books from spines and scanning pages.
What kinds of books is Amazon targeting?
Amazon is targeting older books with lower monetary value—such as books never translated from lesser-used languages or books that were never popular enough to be widely distributed. Workers were trained to scan barcodes or ISBNs before scanning books, suggesting Amazon is methodically working through ISBN lists to ensure high volumes of unique works in its training data.
Do other AI companies do this?
Rivals Anthropic and xAI have publicly stated they are not training on rare or antique books. Amazon declined to comment specifically on its AI training practices.
Ars Technica AIRead Original Article

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Next articleAI agents fail twice as often when given safeguard layers, survey finds

The AI news that matters, in one minute each morning.

Sign up free