AIToday

Pirated books power major AI models, analysis reveals

Hacker News1h agoSend on LINE
Pirated books power major AI models, analysis reveals

Key takeaway

A researcher has confirmed that major AI systems including Meta's LLaMA were trained on upwards of 170,000 pirated books sourced from Bibliotik, a BitTorrent-based collection. The dataset, called Books3, was also used by Bloomberg and EleutherAI and contains work by well-known authors such as Margaret Atwood and Stephen King. While AI companies argue this falls under 'fair use,' legal experts say the law remains unsettled on whether training AI on unauthorized material violates copyright, making this discovery central to ongoing lawsuits against OpenAI and Meta.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    A researcher obtained and analyzed a dataset called Books3 used to train Meta's LLaMA, Bloomberg's BloombergGPT, and EleutherAI's GPT-J. The dataset contains upwards of 170,000 books, the majority published in the past 20 years, sourced from Bibliotik, a pirated-book collection. Named authors whose work appears include Sarah Silverman, Richard Kadrey, Christopher Golden, Michael Pollan, James Patterson, Stephen King, Margaret Atwood (33 books), and Jonathan Franzen (7 books).

  • Why it matters

    The discovery confirms allegations in a lawsuit filed by authors against Meta and OpenAI—that generative AI systems are being trained on copyrighted material without permission. While tech companies argue this constitutes 'fair use,' legal experts say the law is 'unsettled' on whether unauthorized source material strengthens a claim against them. The issue reflects a clash between tech's open-source ethos and publishing's reliance on copyright protection.

  • What to watch

    Books3 was hidden in plain sight for more than two and a half years before Hugging Face's download link stopped working around the time Books3 was mentioned in summer lawsuits. EleutherAI's executive director stated the organization is 'in the process of creating a version of the Pile that exclusively contains documents licensed for that use,' and Bloomberg confirmed it will not include Books3 in future versions of BloombergGPT.

In Depth

In September 2023, a researcher and computer programmer obtained a direct download of 'the Pile,' a massive cache of training text created by EleutherAI. Within that dataset lay Books3, which had been used to train multiple generative AI systems. The researcher then wrote a series of programs to isolate and analyze the Books3 entries, extracting ISBNs from each line and cross-referencing them against an online book database. This process identified more than 170,000 books, the majority published in the past 20 years.

The Books3 dataset was created by Shawn Presser, an independent developer, who downloaded a copy of Bibliotik—a collection of pirated books shared via BitTorrent—from The-Eye.eu and used a program written by hacktivist Aaron Swartz more than a decade earlier to convert the books from ePub format to plain text. Presser announced that Books3 was 'all of Bibliotik.' According to the researcher's analysis, the collection spans publishers of all sizes: more than 30,000 titles from Penguin Random House and its imprints, 14,000 from HarperCollins, 7,000 from Macmillan, 1,800 from Oxford University Press, and 600 from Verso. The breakdown is roughly one-third fiction and two-thirds nonfiction. Named authors include Sarah Silverman, Richard Kadrey, and Christopher Golden—who had filed a lawsuit against Meta alleging copyright violation—as well as Michael Pollan, Rebecca Solnit, Jon Krakauer, James Patterson, Stephen King, George Saunders, Zadie Smith, Junot Díaz, Elena Ferrante, Rachel Cusk, Margaret Atwood (33 books), Jonathan Franzen (7 books), Jennifer Egan (5 books), Haruki Murakami (at least 9 books), bell hooks (9 books), David Grann (5 books), and others. The dataset also includes 102 pulp novels by L. Ron Hubbard and 90 books by John F. MacArthur.

Books3 was used to train Meta's LLaMA, a large language model similar to OpenAI's GPT-4, as well as Bloomberg's initial BloombergGPT and EleutherAI's GPT-J, a popular open-source model. When contacted, a Meta spokesperson declined to comment on the company's use of Books3; a Bloomberg spokesperson confirmed that Books3 was used to train the initial model of BloombergGPT and stated, 'We will not include the Books3 dataset among the data sources used to train future versions of BloombergGPT'; and Stella Biderman, EleutherAI's executive director, did not dispute that the company used Books3 in GPT-J's training data. Biderman also emailed a statement that EleutherAI is 'currently in the process of creating a version of the Pile that exclusively contains documents licensed for that use.'

Presser justified his creation of Books3 by explaining that he feared monopolistic control of generative AI by wealthy corporations and wanted to give independent developers access to 'OpenAI-grade training data.' He told the researcher by telephone that he sympathizes with authors' concerns but believes the greater danger is a monopoly: 'It would be better if it wasn't necessary to have something like Books3. But the alternative is that, without Books3, only OpenAI can do what they're doing.' The article notes that although some titles in Books3 lack copyright-management information, Presser stated he did not knowingly edit the files in this way; the deletions were ostensibly a by-product of file conversion and ebook structure.

The legal landscape remains uncertain. AI companies have argued that training on copyrighted material constitutes 'fair use,' a doctrine that permits use of copyrighted material under certain circumstances to enable parody, quotation, and derivative works. Jason Schultz, director of the Technology Law and Policy Clinic at NYU, said this argument is 'strong' but acknowledged that unauthorized acquisition of source material can damage a fair-use claim. 'If the source is unauthorized, that can be a factor,' he said, but added that 'if they had no idea where the books came from, then I think it's less of a factor.' Rebecca Tushnet, a law professor at Harvard, echoed these points and noted that the law was 'unsettled' when it came to fair-use cases involving unauthorized material, with previous cases providing little indication of how a judge might rule in the future. The article frames this as a clash between the open-source software culture—which emerged in the 1980s when Richard Stallman created 'copyleft' licensing to enable free sharing and modification of code—and the publishing industry's reliance on copyright protection to sustain individual authors' livelihoods.

Context & Analysis

The discovery of Books3 exposes a fundamental tension in how generative AI systems are built. While high-quality AI requires high-quality training data—and books provide that in ways most internet text does not—the AI industry has sourced this data through channels that sidestep copyright law entirely. The dataset was created by Shawn Presser, an independent developer, who justified his approach as a way to democratize AI development and prevent monopolies by wealthy corporations like OpenAI. Presser used a program originally written by hacktivist Aaron Swartz to convert pirated books from ePub format to plain text, and explicitly named Books3 as containing 'all of Bibliotik.' Yet despite being 'hidden in plain sight' for over two and a half years, Books3 became a popular training dataset across the AI industry—cited in research papers by Meta and Bloomberg themselves—until legal pressure from copyright lawsuits forced it underground.

The scale of the piracy is striking: the dataset spans fiction and nonfiction from major publishers (Penguin Random House alone contributes over 30,000 titles) and includes between one-third and two-thirds fiction and nonfiction respectively. The authors represented range from literary figures like Margaret Atwood (33 books) and Jonathan Franzen to genre writers like James Patterson, suggesting that no category of published work was excluded. This wholesale copying reflects what the article frames as a clash of cultures: the open-source software community, which has long operated under permissive licensing models, and the publishing industry, which depends on copyright protection to sustain individual authors. While legal experts acknowledge that 'fair use' doctrine could theoretically shelter AI training from copyright claims, the unauthorized acquisition of the source material introduces a legal uncertainty that has prompted even the AI companies themselves to begin distancing themselves from Books3 going forward.

FAQ

How many books are in the Books3 dataset and who do they come from?
Upwards of 170,000 books are in Books3, the majority published in the past 20 years. The dataset was created by Shawn Presser by downloading Bibliotik, a collection of pirated books shared via BitTorrent, and converting them to plain text. The collection includes more than 30,000 titles from Penguin Random House and its imprints, 14,000 from HarperCollins, 7,000 from Macmillan, and thousands more from other publishers.
Which AI systems have used the Books3 dataset?
Books3 was used to train Meta's LLaMA, Bloomberg's initial BloombergGPT model, and EleutherAI's GPT-J. A Meta spokesperson declined to comment, Bloomberg confirmed it will not use Books3 for future versions of BloombergGPT, and EleutherAI's executive director did not dispute the use.
What is the legal status of training AI on copyrighted books without permission?
AI companies argue it constitutes 'fair use,' which permits copyrighted material to be used under certain circumstances. However, according to legal experts quoted in the article, the law is 'unsettled' on this question. Whether the unauthorized source of the material damages a fair-use claim depends partly on whether the AI companies knew where the books came from, and previous cases provide little indication of how a judge might rule.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime