openai
Microsoft disclaims its own director's "astonishing theft"
Promtime
openaiOne Microsoft internal document about AI training carried a cartoon of large language models destroying their own supply chain. Another, from Microsoft director of applied science Brent Hecht, called scraping news for AI training "an astonishing theft of unprecedented proportions," according to Ars Technica, which read both in filings unsealed Thursday.
At a glance
- The material surfaced in a summary judgment motion from news plaintiffs led by The New York Times, after years in which Microsoft and OpenAI fought to keep those internal documents marked confidential.
- Microsoft's own data, cited in the motion, shows click-through rates falling 83–93 percent for some news plaintiffs and 51–94 percent for others, which the news groups treat as evidence of substitution.
- Microsoft says Hecht's writing reflects one employee's individual perspective and is not legal analysis, and that Satya Nadella's testimony offered observations rather than conclusions about the copyright questions before the court.
If you have not been following the case: according to Wikipedia's account of it, The New York Times sued Microsoft and OpenAI in December 2023, saying OpenAI trained on its content without authorization and that the products can reproduce portions of its articles, while OpenAI called the training fair use and denied being a market substitute. In 2025 Judge Sidney H. Stein denied most of OpenAI's motions to dismiss, and the majority of the claims went forward.
Hecht called it perhaps the largest theft of labor in human history
In documents quoted in the motion, Hecht repeatedly described scraping news for AI training as "an astonishing theft of unprecedented proportions" and perhaps the "largest theft of labor in human history," news organizations said. Another of his documents said the plan to widely scrape news made "a complete mockery of the idea of 'fair use,'" and acknowledged that almost no one intended their work to be used this way, nor are they compensated for its use.
A Microsoft spokesperson defended the company's AI products as a transformative fair use that does not substitute for news sites, and said Hecht's documents "reflect one employee's individual perspective, are not a legal analysis, and do not represent the company's views."
Steven Lieberman, counsel for the New York Daily News and seven of its sister papers, told Ars Technica that the evidence shows OpenAI and Microsoft knew what they were doing was wrong, and that the defendants insisted on confidentiality throughout. "Well, now the cat is out of the bag," he said.
Microsoft logged click-through drops of 83–93 percent for some plaintiffs and 51–94 percent for others
Microsoft's own records, as described in the motion, show click-through rates to some news plaintiffs dropping 83–93 percent, and 51–94 percent for others. ChatGPT head Nick Turley wrote internally that publishers faced an "existential threat" from commercial products trained on news content that can be used to substitute for news providers. One Microsoft document put the economics plainly:
It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its "content supply chain."
Another Microsoft document described a "doom loop" that "will hurt the performance of our models and the entire web at the same time," news groups said, and agreed there is a "real risk" that generative AI could "significantly disrupt" the employment of the very people who generated the data the models were trained on.
Nadella testified that AI platforms hand over the information instead of the source
Under oath, Nadella testified that AI companies should not be violating news sites' terms of use by dodging paywalls. He also described the chatbot as taking clicks from news sites by "giving you the information right there on the website on the AI platform versus needing to go to the underlying source."
Microsoft's spokesperson said that testimony touched on "broad principles and changes underway in how people find and consume information," and that such observations "should not be confused with conclusions about copyright questions before the Court."
At OpenAI, a software engineer wrote that "no matter how prominently we show the links, users won't click." Turley agreed there is "no good reason to click" when the chatbot supplies the information, called chatbots "largely substitutive, period," and predicted they "will get more and more substitutive as they get better."
News groups say OpenAI used a Times dataset of 1.8 million articles
Internal messages show OpenAI staffer Nick Ryder telling president Greg Brockman that "a hack" had been found for OpenAI crawlers "to get around" the NYT paywall. Brockman replied, "Ah, nice."
The motion also accuses Microsoft of violating industry norms by selling a dataset it had purchased for Bing as training data for OpenAI, without consulting news groups whose consent had covered basic search engine crawling. OpenAI, news groups allege, obtained a NYT dataset of 1.8 million articles from a third party bound by an agreement that the data would not be used commercially, and its employees knew it "would not be appropriate" to train a model on it.
The news groups argue the problem runs past these two firms: each company is better off taking content free while others pay, a prisoners' dilemma they say a fair use ruling would settle. They point to Google's AI Overviews, introduced after ChatGPT's launch, absorbing still more traffic that used to reach news sites.
How did news groups get chatbots to reproduce their articles?
Early tests leaned on asking a chatbot "what's the next line?" over and over. The motion shows broader methods: requests for summaries that came back as long excerpts, requests for key bullet points, requests to pick any article off a site's homepage, and, particularly successful, prompts asking the chatbot to "rate the bias" of an article.
The pattern is ordinary once you look at it. Any task that needs the article's text as raw material can push that text into the answer, whether the model carries it from training or pulls it in to ground a response. It is like asking a clerk for an opinion on a document and getting the document read back to you.
News plaintiffs also allege that instead of preventing such outputs, the firms built a filter that Hecht suggested could be perceived as an "accidental cover up," because it would result in "people who have a right over the content having less visibility into what was used for training."
Read carefully, all of this is one side's selection: the quotes come from a motion the news plaintiffs wrote, Microsoft disputes their framing, and no court has ruled on any of it. The click-through figures arrive without a stated time window or baseline. In our view the filter is the oddest item in the pile, since by Hecht's own description it left the people who own the content seeing less of what was used to train on it.
What the verbatim-overlap ruling covers
The news plaintiffs asked the court to rule now only on infringed articles whose outputs "demonstrate extensive verbatim overlap," saying legal concerns with the other articles will be raised at trial. They say they are ready for that trial because of the evidence of substitution they have gathered. Nothing in the unsealed motion gives a date for a ruling or for the trial itself.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
