AIToday
Large Language ModelsAI Safety & AlignmentFortune AIPublished: Jul 31, 2026, 19:00 JST5 min read

AI companies quietly buying thousands of used books to train models

AI companies quietly buying thousands of used books to train models

Key takeaway

  • A Dutch antiquarian bookseller received an unusual bulk book order for over 3,000 titles that he initially dismissed as spam, but investigation revealed it was connected to AI training.

  • The discovery exposed a hidden supply chain where AI companies are quietly acquiring large quantities of secondhand books—including through destructive scanning, where spines are cut and pages fed through high-speed scanners before the originals are discarded.

  • A federal judge previously ruled this practice constituted fair use under copyright law when Anthropic used legally purchased books to train its models.

3 Key Points

  1. What happened

    A Dutch bookseller received an email from a company called 2077AI requesting 3,001 book titles—mostly academic works published between 2020 and 2021—for shipment to China, but dismissed it as spam until a journalist revealed the request's connection to artificial intelligence. Similar bulk book orders have surfaced in Germany and Switzerland, suggesting AI companies are building physical supply chains to source books at scale.

  2. Why it matters

    Court records showed Anthropic purchased millions of physical books, cut their spines, scanned them, and discarded the originals to train AI models—a practice a federal judge ruled was fair use under copyright law. The hidden procurement networks reveal how AI companies are systematically acquiring secondhand books outside public view, raising questions about the sourcing of training data for large language models.

  3. What to watch

    Earlier this year, ISBNdb advertised services to help AI labs source bulk book purchases ranging from 1,000 to 1 million copies, though the company later said the service was never launched and the webpages reflected only an exploratory concept. The extent of ongoing book acquisition by AI companies remains unclear, as 2077AI did not respond to requests for comment.

In Depth

Read the full story

On a summer day in Haarlem, Netherlands, antiquarian bookseller Pieter de Vries received an email from a woman named Natalia, identifying herself as representing a company called "2077AI." The message requested what she called a "fairly large order" and included a spreadsheet with 3,001 ISBNs and instructions to match titles, prepare quotes, and estimate shipping costs to China. De Vries barely glanced at it, dismissing the message as spam or phishing without reading past the first lines. It was only weeks later, when a Dutch journalist investigating the procurement contacted him, that de Vries learned the request was connected to artificial intelligence. "I was shocked!" he told Fortune.

The 3,001 titles in the spreadsheet were predominantly academic works published between 2020 and 2021 by major publishers including Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, spanning subjects from business and education to engineering, public policy and medicine. De Vries was not alone in receiving the request—several other Dutch antiquarian booksellers got the same message and similarly dismissed it as a scam. For de Vries, who specializes in old and rare books and rarely receives bulk orders, the unusual nature of the request reinforced his skepticism. "These books are expensive and exclusive," he said. "So I do not do bulk trade."

The inquiry provided a rare window into how AI companies are sourcing training material beyond the open internet. Last summer, court records revealed that Anthropic had purchased millions of physical books, removed their bindings, scanned them, and discarded the originals to build a searchable digital library for training its AI models in what internal documents called "Project Panama," according to The Washington Post. The process, known as "destructive scanning," involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is thrown away. A federal judge later ruled that using legally purchased books to train AI models constituted fair use under copyright law, and the case settled. Anthropic stated that Claude models are trained on a mix of publicly available web data, commercially acquired datasets and internally generated data, and that it purchases books through regular commercial markets.

Earlier this year, 404 Media reported that ISBNdb—a company known for maintaining metadata tied to International Standard Book Numbers, the unique numerical identifiers found on most book barcodes—had begun advertising services to help AI labs source bulk printed book purchases ranging from 1,000 to 1 million books. According to the report, the company's website said the service was tailored to "LLM training needs" and "delivered at the scale AI demands." ISBNdb later told Fortune the service was never launched and that the webpages reflected only an exploratory concept. Similar reports have emerged in Germany and Switzerland, where local media reported that secondhand booksellers received unusual bulk orders for highly specialized books inconsistent with traditional collecting or resale. Zoom Books, a Canadian company named in some reports, denied the allegations and said the purchases were part of its regular recycling and trading model. Fortune contacted 2077AI for comment about the purpose of the procurement request and whether the books were intended for AI training, but the company did not respond.

Context & Analysis

The email to Dutch bookseller Pieter de Vries offers a concrete example of how AI companies are building physical supply chains to acquire training data. While the public debate around AI training data has focused on web scraping and copyright disputes over digital libraries, this procurement request reveals a parallel, largely invisible operation: the systematic acquisition of physical books through secondhand markets. The request itself—3,001 academic titles shipped to China—bears the hallmarks of deliberate sourcing rather than casual purchasing, yet was designed to look innocuous enough that multiple booksellers dismissed it as spam.

The connection to Anthropic's "Project Panama" provides crucial context. Court records revealed that the company purchased millions of physical books, destructively scanned them, and disposed of the originals—a process that federal courts later sanctioned as fair use. This legal outcome may have emboldened broader procurement efforts across the AI industry. The discovery that ISBNdb, a metadata company, advertised bulk book sourcing services for "LLM training needs" at scales of "1,000 to 1 million books" suggests that infrastructure for large-scale acquisition is being built, even if ISBNdb later claimed the service was never operational. Similar orders reported in Germany and Switzerland indicate this is not an isolated incident but part of a coordinated, international supply strategy.

FAQ

What books were requested in the order to the Dutch bookseller?
The order listed 3,001 titles, mostly published between 2020 and 2021 by academic publishers including Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, covering subjects ranging from business and education to engineering, public policy and medicine.
How did the court rule on AI companies using purchased books for training?
A federal judge ruled that using legally purchased books to train AI models constituted fair use under copyright law, and the case settled after Anthropic resolved separate claims involving the downloading of books from online libraries.
What is destructive scanning?
Destructive scanning involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded, a process Anthropic used when building its searchable digital library for AI training.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Related Articles

Next articleMicrosoft shifts AI strategy to multi-model platform play

The AI news that matters, in one minute each morning.

Sign up free