AIToday

Anthropic pays $1.5B to settle book piracy claims—AI training on legal content upheld as fair use

THE DECODER1h ago
Anthropic pays $1.5B to settle book piracy claims—AI training on legal content upheld as fair use

Key takeaway

Anthropic has agreed to pay book authors $1.5 billion(約2400億円) in a federal court-approved settlement for downloading works from piracy databases between 2021 and 2022—the largest copyright settlement in class action history. However, a judge previously ruled that training AI on legally obtained books is fair use, a distinction that may protect AI labs that scraped web content without owner consent, their primary source of training data. The settlement covers the piracy act itself, not AI training, leaving the broader question of whether mass web scraping without consent is legal still unresolved.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Anthropic agreed to pay book authors $1.5 billion(約2400億円) to settle a copyright lawsuit over downloading roughly 482,460 works from piracy databases LibGen and PiLiMi between 2021 and 2022. A federal court in San Francisco approved the settlement, with authors claiming about $3,000 each—four times the statutory minimum. Anthropic must destroy the pirated files.

  • Why it matters

    The settlement is the largest copyright settlement in class action history, but it covers the piracy itself, not AI training. Judge Alsup previously ruled that training AI on legally obtained books is "transformative—spectacularly so" and falls under fair use. This distinction means AI labs that trained on web-scraped content without website owners' consent may face reduced legal exposure, since the core question of whether that scraping was lawful remains unsettled.

  • What to watch

    Authors retain claims over AI outputs that reproduce their original works and over Anthropic's future conduct. The fair use debate over mass scraping of internet content without consent is likely far from over, so the legal landscape for AI training data remains in flux.

In Depth

In a landmark settlement approved by a federal court in San Francisco, Anthropic agreed to pay book authors $1.5 billion(約2400億円) to resolve a copyright lawsuit stemming from its download of works from the piracy databases LibGen and PiLiMi. Between 2021 and 2022, Anthropic obtained roughly 482,460 listed works from these unlicensed sources. Of that collection, 91.3 percent were claimed by authors in the settlement class, resulting in a payout of about $3,000 per work—four times the statutory minimum. The settlement is the largest copyright settlement in class action history. As part of the agreement, Anthropic must destroy the pirated files and faces ongoing liability: authors retain claims over AI outputs that reproduce their original works and over Anthropic's future conduct.

The settlement, however, tells only part of the legal story. Judge Alsup had previously issued a ruling that stands as a shield for AI training practice more broadly. He determined that training AI on legally obtained books is "transformative—spectacularly so" and constitutes fair use under copyright law. This ruling means that the act of training an AI model on lawful data does not, in itself, infringe copyright. The settlement thus addresses the piracy—the unlicensed acquisition of the books—not the use of those books in AI training. This distinction is significant because it leaves unresolved a broader legal question: whether the mass scraping of internet content without owners' consent, the primary source of training data for most AI labs, counts as lawful acquisition in the first place. The ruling effectively hands AI labs a major victory on the fair use question while leaving the acquisition question open. Whether the scraping itself is legal remains unsettled, and the article suggests that "the fair use debate is likely far from over," implying that future litigation could test the boundaries further and potentially reshape the legal terrain for data acquisition practices.

Context & Analysis

The settlement represents a watershed moment framed very differently depending on where you stand. For authors and publishers, it is vindication: $1.5 billion(約2400億円) is the largest copyright settlement in class action history, and the per-work payout of roughly $3,000 is four times the statutory minimum, signaling judicial seriousness about author harm. For Anthropic and other AI labs, however, the ruling contains a protective aperture. Judge Alsup's prior determination that AI training on legally obtained books is "transformative—spectacularly so" and constitutes fair use creates a legal scaffold that insulates many training practices. The settlement itself does not target training; it targets piracy—the act of downloading unlicensed works. This distinction is material. If training on lawful data is fair use, then the bottleneck for AI labs shifts from the training step itself to the question of whether their data acquisition was lawful. That question, the article makes clear, remains open. Most AI labs have sourced training data by scraping the web and internet archives without owner consent. The ruling does not resolve whether that scraping is legal; it only clarifies that once the data is lawfully in hand, using it for AI training is likely permissible. The fair use debate is "likely far from over," the article notes, suggesting that litigation over the legality of the scraping itself could emerge separately and test the boundaries further.

FAQ

What exactly did Anthropic do that led to the settlement?
Anthropic downloaded roughly 482,460 books from piracy databases LibGen and PiLiMi between 2021 and 2022. The settlement requires Anthropic to pay $1.5 billion(約2400億円) to authors and destroy the pirated files.
Does this ruling mean AI training on copyrighted content is illegal?
No. Judge Alsup previously ruled that training AI on legally obtained books is "transformative—spectacularly so" and falls under fair use. The settlement covers piracy, not AI training itself. Whether mass scraping of internet content without authors' consent counts as legal acquisition remains an open question.

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No discussion yet for this article

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →