OpenAI's 2019 note: "We trained GPT-3 on pirated stuff"

Among the unsealed material in the Authors Guild case is a tweet quoted in testimony: if George R.R. Martin dies early, "GPT can write the next [S]ong of [I]ce and [F]ire book." The same brief, unsealed less than two weeks after both sides moved for summary judgment, quotes OpenAI staff on pirated training data and on authors losing work, as Publishers Weekly reports.
At a glance
- OpenAI hired Tarun Gogineni in 2022 to improve the writing quality of its models; the documents show he described his research mission as building a machine that would supplant human authors.
- OpenAI policy director Jack Clark testified in May 2020 that genre fiction authors would worry about substitution on Amazon, and that the company would likely ignore artists' concerns and release anyway.
- Microsoft knew about OpenAI's LibGen use as early as April 2019, the brief states, and in 2022 an effort called Project Clear removed the LibGen files from OpenAI's systems.
If you missed the earlier rounds, book authors have been suing AI labs over shadow libraries for two years. According to NPR, Anthropic agreed in August 2025 to pay 1.5 billion dollars to settle a class action over training Claude on pirated books, about 3,000 dollars a book for 482,460 works, and to destroy the files it had downloaded.
An OpenAI note from August 2019 says GPT-3 was trained on pirated material
The bluntest line in the unsealed material is also the shortest. An OpenAI note from August 2019, quoted in the class plaintiffs' brief, carries no legal framing, no qualifier and no euphemism for where the books came from. It reads in full:
We trained GPT-3 on pirated stuff! No sharing that!
Jack Clark, OpenAI's policy director, testified in May 2020. "The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon," he said. He also acknowledged that there would be a point where a bunch of artists express worry about the work and, in his words, "we'll likely ignore their concerns and release anyway."
Tarun Gogineni called author complaints "acceptable economic disruption"
OpenAI hired Tarun Gogineni in 2022 to lead work on the writing quality of its models. The documents show he described his "research mission" as building a machine that would supplant human authors. The brief states he knew about the complaints that the "datasets are stolen" and that authors were losing work to AI-generated competition.
He did not find those complaints "all that sympathetic," according to the brief, and treated them as "acceptable economic disruption." He also wrote that the world would soon see "the death of the reader" as "machines creat[ed] slop for more machines."
Publishers Weekly reports that the unsealed comments show executives knew they were training on copyrighted books illegally and that AI-generated books would likely displace human-authored books in the market, and that some publishing industry employees will lose their jobs. The Authors Guild highlighted several of the now public remarks by OpenAI executives.
Microsoft knew about the LibGen use as early as April 2019
That is what the brief states. In 2022, through an effort called Project Clear, OpenAI deleted its LibGen files. The memo behind it came from OpenAI VP of research Bob McGraw, who wrote to colleagues: "Given how much OpenAI is in the news, now is the right time to excise Libgen from our systems and storage. What would be involved in that?"
The quotes come from two filings: the class plaintiffs' memorandum in support of partial summary judgment, Docket 1982, and their corrected Rule 56.1 statement of undisputed material facts, Docket 1987. They landed days after unsealed files in the New York Times suit against Microsoft and OpenAI showed that building large language models on millions of news articles posed an "existential threat" to newspaper publishers.
According to The Wall Street Journal, Microsoft's director of applied science Brent Hecht wrote soon after the Times sued that the companies had started a "doom loop" that will hurt the performance of their models and the entire web at once, and that an end-product threatening the economic foundations of its essential suppliers is highly unusual.
Why does deleting the LibGen files matter?
Because the copying and the training have been judged as separate acts. Judge Alsup's June 2025 ruling drew a line between fair-use training on legally acquired, digitized books and downloading and permanently retaining pirated copies; for the pirated copies specifically, he found that every factor points against fair use.
Think of two libraries. One scans the books it bought and keeps the scans on a server; the other keeps a stolen crate in the basement and reads from that. The scanning and the basement are judged separately, even when the reading looks the same from outside.
According to Whitecase's account of two parallel rulings two days apart in June 2025, Judge Chhabria found Meta's training of Llama on shadow libraries highly transformative, and both judges said different economic evidence could have changed the outcome.
All of this comes from the plaintiffs' own filings, which quote the lines that help their case and give them without the conversations around them. In our view, deleting the LibGen files in 2022 is an odd remedy for material the brief says Microsoft knew about in April 2019.
What partial summary judgment decides
The class plaintiffs are asking for partial summary judgment, so the next ruling will not end the case; it will set which pieces go to trial and which are already settled. The open questions are the ones the documents circle: whether keeping pirated files is separable from training on them, and whether displacement of human-authored books can be shown. No hearing or ruling date appears in the filings quoted here.
Related stories
- Microsoft disclaims its own director's "astonishing theft"
- Both sides seek a ruling without trial in AI book case
- Oxford's Bodleian Library texts ended up training OpenAI
- Two more newspapers take OpenAI to court
- 30 more lawsuits over Tumbler Ridge school shooting
- US government sides with OpenAI in NYT lawsuit
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
