AIToday
Large Language ModelsAI Safety & AlignmentArs Technica AIPublished: Aug 13, 2026, 01:00 JST5 min read

AI firms quietly buying rare books, destroying them to train models

AI firms quietly buying rare books, destroying them to train models

Key takeaway

  • AI companies are bulk-buying and destroying rare books to train their models, sparking alarm among booksellers and book lovers.

  • A lawsuit in summer 2025 exposed Anthropic's destruction of millions of print books, prompting rare booksellers to flag suspicious large orders as red flags for AI involvement.

  • While non-destructive scanning methods exist—such as the Internet Archive's careful hand-scanning process and Google's patented technology—cost and speed pressures may lead AI firms to continue the cheaper practice of destroying originals.

3 Key Points

  1. What happened

    A lawsuit in summer 2025 revealed that Anthropic destroyed millions of print books to train its AI models. Since then, rare booksellers have reported suspicious bulk orders—including an Irish bookstore that flagged an order for 5,000 obscure titles with no price haggling, a now-obvious red flag for AI involvement.

  2. Why it matters

    Rare books cannot be replaced once destroyed. AI firms can legally purchase and destroy books they own, but the practice threatens the survival of hard-to-find titles and undermines authors' intellectual property rights. Some booksellers now refuse orders suspected of destined for AI training, fearing they are "signing their own death warrant" by enabling the loss of their inventory.

  3. What to watch

    Non-destructive alternatives exist—Google patented a non-destructive book-scanning method in 2009, and the Internet Archive uses labor-intensive hand-scanning to preserve rare books physically while digitizing them. OpenAI and Microsoft are partnering with Harvard librarians to train AI on about 1 million public-domain books from the 15th century onward. Whether major AI firms adopt preservation-focused methods or continue cutting corners will determine whether rare collections survive.

In Depth

Read the full story

In summer 2025, a lawsuit publicly revealed that Anthropic had destroyed millions of print books as part of its AI model training process. The disclosure sparked immediate backlash from book lovers and the broader publishing community, who fear that AI companies' hunger for high-quality long-form text will consume rare and irreplaceable collections. The core problem is straightforward: the cheapest and fastest way to digitize books is to cut off their spines, feed pages into scanning machines, and discard the originals. This method allows AI firms to process millions of titles quickly—a critical advantage in the race to advance their models—but it means rare books are lost forever once destroyed.

In response to the outcry, alternative technologies have been highlighted. Google patented a non-destructive scanning method in 2009, but studies have found that page curvature can distort text and pages can be missed, and workers may move too quickly, introducing errors like hands obscuring pages. The Internet Archive has taken a different path entirely, choosing labor over speed. A 2021 blog post detailed the work of Eliza Zhang, a book scanner at the Archive since 2010, who hand-scans rare books one page at a time. She raises the scanner glass with a foot pedal, adjusts cameras, and ensures readability before moving to the next page. She also documents fold-outs and inserts to prevent losing bonus materials. Her proprietary software halts the process if pages are skipped or images blur, prompting rescans. Zhang had scanned more than 3 million pages, 14,000 foldouts, and 18,000 items (mostly books) with a goal of zero errors. The Internet Archive rejected automated scanners with vacuum-powered page-turning arms because they damaged brittle books and rare volumes—the very collections librarians asked them to preserve.

Meanwhile, rare booksellers have begun flagging suspicious bulk orders. Last week, Kennys Bookshop in Ireland reported a "bananas" order for 5,000 obscure titles, with no attempt to negotiate price—a red flag that signals AI involvement. Unlike typical library or university purchases, these orders consist of widely varying, unrelated selections bought in massive batches at full price. Tomás Kenny of Kennys Bookshop told the Irish Times that AI is "terrifying our industry" and pledged to refuse orders suspected of destined for AI training to protect authors' rights and readers' access to hard-to-find titles. Booksellers worry that short-term revenue from such orders amounts to "signing their own death warrant," because once rare books are destroyed to train AI, they vanish from the supply chain forever.

Not all AI firms have taken a destructive path. OpenAI and Microsoft are partnering with Harvard librarians to train AI models on about 1 million public-domain books dating back to the 15th century. Elon Musk posted on X that he asked the xAI team to preserve rare books and scan them the hard way instead of cutting spines, though critics note this does not constitute a firm commitment to avoid destroying books entirely. Anthropic has denied destroying rare books and has provided no public accounting of its practices. The company reportedly warned its clients via ISBNdb—a book-database company advertising bulk sourcing services to AI firms—that the optics of book destruction were bad, and adopted the codename "Project Panama" to hide the effort from public view. Copyright law permits companies to destroy copies of books they purchase, but public backlash and the viability of preservation-focused methods suggest that the industry's long-term reputation may depend on choosing slower, more careful approaches over cost-cutting shortcuts.

Context & Analysis

The tension between AI training efficiency and book preservation reflects a deeper clash between speed-to-market and stewardship. A lawsuit in summer 2025 exposed Anthropic's destruction of millions of print books, and the backlash has forced the company to adopt a codename—"Project Panama"—for its destructive scanning efforts in an attempt to hide the practice from public scrutiny. The discovery has mobilized rare booksellers, who now recognize suspicious bulk orders as markers of AI involvement: large batches of unrelated titles purchased without price negotiation stand out sharply from typical library or university acquisitions. The industry's concern is not merely sentimental—it reflects a real threat to the survival of irreplaceable collections and an erosion of authors' intellectual property rights. Yet the economics of AI training create powerful incentives to cut corners. Non-destructive alternatives do exist: Google's 2009 patent, though imperfect, offers a blueprint, and the Internet Archive's labor-intensive hand-scanning—developed over decades and perfected by scanners like Eliza Zhang, who has processed more than 3 million pages and 18,000 items—demonstrates that preservation and digitization can coexist. The question now is whether public backlash and the commitment of firms like OpenAI and Microsoft to preserve public-domain collections will be enough to shift the broader industry toward preservation-focused methods, or whether cost pressures will continue to drive destruction.

FAQ

How are AI companies destroying books to train their models?
The cheapest and fastest method is to cut off book spines and feed pages into machines for scanning, then discard them. This avoids the time and labor investment required by non-destructive methods like hand-scanning used by the Internet Archive.
What non-destructive scanning methods exist?
Google patented a non-destructive book-scanning technology in 2009, though studies show page curves can distort text and pages may be missed. The Internet Archive uses careful hand-scanning—workers use foot pedals to raise scanner glass, adjust cameras, and ensure readability page by page—a method the Internet Archive says works best for fragile and rare books, with proprietary software catching errors and prompting rescans.
Are any AI firms committing to preserve rare books?
OpenAI and Microsoft are working with Harvard librarians to train AI models on about 1 million public-domain books dating back to the 15th century. Elon Musk said he asked the xAI team to scan rare books the hard way instead of destroying them, though critics note he did not commit xAI to avoiding book destruction entirely.
Ars Technica AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articlePixel 11 Gets AI Photo-Capture Mode, 120X Zoom, Faster Night Shots

The AI news that matters, in one minute each morning.

Sign up free