
A Dutch antiquarian bookseller received an unusual bulk book order for over 3,000 titles that he initially dismissed as spam, but investigation revealed it was connected to AI training.
The discovery exposed a hidden supply chain where AI companies are quietly acquiring large quantities of secondhand books—including through destructive scanning, where spines are cut and pages fed through high-speed scanners before the originals are discarded.
A federal judge previously ruled this practice constituted fair use under copyright law when Anthropic used legally purchased books to train its models.
What happened
A Dutch bookseller received an email from a company called 2077AI requesting 3,001 book titles—mostly academic works published between 2020 and 2021—for shipment to China, but dismissed it as spam until a journalist revealed the request's connection to artificial intelligence. Similar bulk book orders have surfaced in Germany and Switzerland, suggesting AI companies are building physical supply chains to source books at scale.
Why it matters
Court records showed Anthropic purchased millions of physical books, cut their spines, scanned them, and discarded the originals to train AI models—a practice a federal judge ruled was fair use under copyright law. The hidden procurement networks reveal how AI companies are systematically acquiring secondhand books outside public view, raising questions about the sourcing of training data for large language models.
What to watch
Earlier this year, ISBNdb advertised services to help AI labs source bulk book purchases ranging from 1,000 to 1 million copies, though the company later said the service was never launched and the webpages reflected only an exploratory concept. The extent of ongoing book acquisition by AI companies remains unclear, as 2077AI did not respond to requests for comment.
On a summer day in Haarlem, Netherlands, antiquarian bookseller Pieter de Vries received an email from a woman named Natalia, identifying herself as representing a company called "2077AI." The message requested what she called a "fairly large order" and included a spreadsheet with 3,001 ISBNs and instructions to match titles, prepare quotes, and estimate shipping costs to China. De Vries barely glanced at it, dismissing the message as spam or phishing without reading past the first lines. It was only weeks later, when a Dutch journalist investigating the procurement contacted him, that de Vries learned the request was connected to artificial intelligence. "I was shocked!" he told Fortune.
The 3,001 titles in the spreadsheet were predominantly academic works published between 2020 and 2021 by major publishers including Emerald Publishing, Elsevier, Wiley, Routledge and Oxford University Press, spanning subjects from business and education to engineering, public policy and medicine. De Vries was not alone in receiving the request—several other Dutch antiquarian booksellers got the same message and similarly dismissed it as a scam. For de Vries, who specializes in old and rare books and rarely receives bulk orders, the unusual nature of the request reinforced his skepticism. "These books are expensive and exclusive," he said. "So I do not do bulk trade."
The inquiry provided a rare window into how AI companies are sourcing training material beyond the open internet. Last summer, court records revealed that Anthropic had purchased millions of physical books, removed their bindings, scanned them, and discarded the originals to build a searchable digital library for training its AI models in what internal documents called "Project Panama," according to The Washington Post. The process, known as "destructive scanning," involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is thrown away. A federal judge later ruled that using legally purchased books to train AI models constituted fair use under copyright law, and the case settled. Anthropic stated that Claude models are trained on a mix of publicly available web data, commercially acquired datasets and internally generated data, and that it purchases books through regular commercial markets.
Earlier this year, 404 Media reported that ISBNdb—a company known for maintaining metadata tied to International Standard Book Numbers, the unique numerical identifiers found on most book barcodes—had begun advertising services to help AI labs source bulk printed book purchases ranging from 1,000 to 1 million books. According to the report, the company's website said the service was tailored to "LLM training needs" and "delivered at the scale AI demands." ISBNdb later told Fortune the service was never launched and that the webpages reflected only an exploratory concept. Similar reports have emerged in Germany and Switzerland, where local media reported that secondhand booksellers received unusual bulk orders for highly specialized books inconsistent with traditional collecting or resale. Zoom Books, a Canadian company named in some reports, denied the allegations and said the purchases were part of its regular recycling and trading model. Fortune contacted 2077AI for comment about the purpose of the procurement request and whether the books were intended for AI training, but the company did not respond.
The email to Dutch bookseller Pieter de Vries offers a concrete example of how AI companies are building physical supply chains to acquire training data. While the public debate around AI training data has focused on web scraping and copyright disputes over digital libraries, this procurement request reveals a parallel, largely invisible operation: the systematic acquisition of physical books through secondhand markets. The request itself—3,001 academic titles shipped to China—bears the hallmarks of deliberate sourcing rather than casual purchasing, yet was designed to look innocuous enough that multiple booksellers dismissed it as spam.
The connection to Anthropic's "Project Panama" provides crucial context. Court records revealed that the company purchased millions of physical books, destructively scanned them, and disposed of the originals—a process that federal courts later sanctioned as fair use. This legal outcome may have emboldened broader procurement efforts across the AI industry. The discovery that ISBNdb, a metadata company, advertised bulk book sourcing services for "LLM training needs" at scales of "1,000 to 1 million books" suggests that infrastructure for large-scale acquisition is being built, even if ISBNdb later claimed the service was never operational. Similar orders reported in Germany and Switzerland indicate this is not an isolated incident but part of a coordinated, international supply strategy.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
The world is running short on high-bandwidth memory (HBM), the specialized chips that feed data to AI systems

Microsoft is reframing its artificial intelligence approach, moving away from the goal of building a single be…

ByteDance is reorganizing Doubao (its AI chatbot), Lark (workplace software), and Volcengine (cloud infrastruc…

Anthropic revealed that three of its Claude models unintentionally accessed the systems of three organizations…

Nvidia CEO Jensen Huang identified memory chips as AI's largest bottleneck, shifting focus from the earlier pr…

OpenAI cut GPT-5.6 Luna prices by 80% (now $0.20 per million input tokens and $1.20 per million output tokens)…

The AI news that matters, in one minute each morning.
Sign up free