
What happened
The Authors Guild's September 17, 2026 filings include messages from OpenAI researcher Tarn Goganeni and then-executive Dario Amodei, and describe a 2022 project that removed LibGen files.
Why it matters
The documents suggest OpenAI's leadership understood the training data could harm authors, which appears to undercut any claim that the risk was unforeseen.
What to watch
The filings are the Authors Guild's allegations in an ongoing lawsuit against OpenAI and Microsoft, so the outcome hinges on how the court weighs these internal messages.
WHO IT HITSLiterary authors and publishers whose copyrighted works may have been used in training data gain new internal-document evidence for their claims. Enterprise legal teams at AI developers may see this as a warning about how internal messages read in litigation.
Summaries like this, in your inbox every morning.
The filings trace a pattern that starts well before the current generative AI boom. In May 2020, Jack Clark, then OpenAI's policy director and later an Anthropic co-founder, wrote internally that the company's work would lead to systems replacing people's labor and that OpenAI would likely ignore artists' concerns and ship anyway. By 2019, according to the documents, OpenAI had already shown Microsoft an early GPT-3 and disclosed that it had used Library Genesis (LibGen), described as one of the world's largest pirate libraries, for training.
The internal debate that followed appears to have focused less on legality than on optics. Sam McCandlish said his worry was how it would look — specifically an article on Hacker News accusing OpenAI of using copyrighted data from a questionable Russian site. Dario Amodei was said to have had some doubts about LibGen as a training set. By summer 2022, OpenAI ran Project Clear, deleting LibGen files; a June 15, 2022 Slack exchange shows research VP Bob McGrew deciding to purge LibGen from systems and storage because OpenAI was in the news so much.
The Authors Guild's CEO, Mary Rasenberger, argues the documents show OpenAI and Microsoft held writers and their work in astonishing disregard, choosing to repeatedly steal books rather than pay for them. The case's outcome may hinge on whether a court reads these messages as deliberate infringement or as the ordinary internal debate of a fast-moving startup, and on how much weight it gives to the executives' own stated expectations about the harm to authors.
For example, today's edition would include:
AI-summarized, only the topics you pick: one digest a day via Email, LINE, or Slack.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
The AI Conference announced its full agenda, with 120+ speakers and an anticipated 5,000 attendees at Pier 48…

Sandhya Venkatachalam, founder of Axiom Partners, said her $52 million fund expects to make 35 investments, wi…

Ryan Greenblatt, chief scientist at Redwood Research, said on Sam Harris's podcast that roughly a 50 to 60 per…

Akamai Technologies signed a seven-year, $11.6 billion cloud infrastructure services (CIS) contract with Anthr…

Anthropic CEO Dario Amodei had dinner with President Trump at the White House on Sunday, after critics circula…

OpenAI paused training its most powerful models and notified "dozens" of governments, universities, and public…
