
Anthropic has agreed to pay book authors $1.5 billion(約2400億円) in a federal court-approved settlement for downloading works from piracy databases between 2021 and 2022—the largest copyright settlement in class action history. However, a judge previously ruled that training AI on legally obtained books is fair use, a distinction that may protect AI labs that scraped web content without owner consent, their primary source of training data. The settlement covers the piracy act itself, not AI training, leaving the broader question of whether mass web scraping without consent is legal still unresolved.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Anthropic agreed to pay book authors $1.5 billion(約2400億円) to settle a copyright lawsuit over downloading roughly 482,460 works from piracy databases LibGen and PiLiMi between 2021 and 2022. A federal court in San Francisco approved the settlement, with authors claiming about $3,000 each—four times the statutory minimum. Anthropic must destroy the pirated files.
Why it matters
The settlement is the largest copyright settlement in class action history, but it covers the piracy itself, not AI training. Judge Alsup previously ruled that training AI on legally obtained books is "transformative—spectacularly so" and falls under fair use. This distinction means AI labs that trained on web-scraped content without website owners' consent may face reduced legal exposure, since the core question of whether that scraping was lawful remains unsettled.
What to watch
Authors retain claims over AI outputs that reproduce their original works and over Anthropic's future conduct. The fair use debate over mass scraping of internet content without consent is likely far from over, so the legal landscape for AI training data remains in flux.
In a landmark settlement approved by a federal court in San Francisco, Anthropic agreed to pay book authors $1.5 billion(約2400億円) to resolve a copyright lawsuit stemming from its download of works from the piracy databases LibGen and PiLiMi. Between 2021 and 2022, Anthropic obtained roughly 482,460 listed works from these unlicensed sources. Of that collection, 91.3 percent were claimed by authors in the settlement class, resulting in a payout of about $3,000 per work—four times the statutory minimum. The settlement is the largest copyright settlement in class action history. As part of the agreement, Anthropic must destroy the pirated files and faces ongoing liability: authors retain claims over AI outputs that reproduce their original works and over Anthropic's future conduct.
The settlement, however, tells only part of the legal story. Judge Alsup had previously issued a ruling that stands as a shield for AI training practice more broadly. He determined that training AI on legally obtained books is "transformative—spectacularly so" and constitutes fair use under copyright law. This ruling means that the act of training an AI model on lawful data does not, in itself, infringe copyright. The settlement thus addresses the piracy—the unlicensed acquisition of the books—not the use of those books in AI training. This distinction is significant because it leaves unresolved a broader legal question: whether the mass scraping of internet content without owners' consent, the primary source of training data for most AI labs, counts as lawful acquisition in the first place. The ruling effectively hands AI labs a major victory on the fair use question while leaving the acquisition question open. Whether the scraping itself is legal remains unsettled, and the article suggests that "the fair use debate is likely far from over," implying that future litigation could test the boundaries further and potentially reshape the legal terrain for data acquisition practices.
The settlement represents a watershed moment framed very differently depending on where you stand. For authors and publishers, it is vindication: $1.5 billion(約2400億円) is the largest copyright settlement in class action history, and the per-work payout of roughly $3,000 is four times the statutory minimum, signaling judicial seriousness about author harm. For Anthropic and other AI labs, however, the ruling contains a protective aperture. Judge Alsup's prior determination that AI training on legally obtained books is "transformative—spectacularly so" and constitutes fair use creates a legal scaffold that insulates many training practices. The settlement itself does not target training; it targets piracy—the act of downloading unlicensed works. This distinction is material. If training on lawful data is fair use, then the bottleneck for AI labs shifts from the training step itself to the question of whether their data acquisition was lawful. That question, the article makes clear, remains open. Most AI labs have sourced training data by scraping the web and internet archives without owner consent. The ruling does not resolve whether that scraping is legal; it only clarifies that once the data is lawfully in hand, using it for AI training is likely permissible. The fair use debate is "likely far from over," the article notes, suggesting that litigation over the legality of the scraping itself could emerge separately and test the boundaries further.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No discussion yet for this article
Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack