Technology
Danish Kapoor
Danish Kapoor

Microsoft executive describes artificial intelligence data use as ‘labor theft’

New documents made public in the copyright lawsuit filed by The New York Times against OpenAI and Microsoft nearly three years ago have brought the debate about companies’ methods of training artificial intelligence models and the use of publishers’ content to the agenda again. According to internal company memos cited in the uncensored dossier submitted to the court by The Times, senior Microsoft employees criticized large-scale content collection methods in extremely harsh terms. The same filing claims that OpenAI executives acknowledged that products like ChatGPT could put serious pressure on news publishers’ business models. The documents also include new details about practices such as accessing paid content, retrieving data from the Bing index, using Common Crawl resources, and removing copyright notices from training data. However, a significant part of this new information does not come directly from the entire internal company documents, but from excerpts from the legal text submitted to the court by The New York Times, and some of the relevant annexes are still not publicly available.

At the center of the case is the question of whether it is legally acceptable to use copyrighted news content without permission in the training of generative artificial intelligence models. OpenAI and Microsoft argue that such use can be considered within the scope of the “fair use” principle in US copyright law. This principle allows the use of copyrighted works without permission from the rights holder under certain conditions, but the evaluation takes into account the nature of the use as well as its impact on the economic market of the original work. The internal statements highlighted by The Times in its latest application also focus specifically on this economic impact topic. It is stated in the source text that the Trump administration recently filed a court opinion defending OpenAI’s use of unlicensed educational data, but since this current political and legal development cannot be independently verified, it is quoted here only as a claim in the source text.

Microsoft documents also bring up the decline in publisher traffic

According to Microsoft data cited in the court filing, the use of Copilot in its “response engine” format resulted in clicks directed to The New York Times domain decreasing by as much as 93 percent in some cases compared to a traditional Bing search. In an internal presentation prepared by Microsoft Applied Sciences Director Brent Hecht in January 2024, this situation is defined as a “doom loop” and it is stated that the weakening of content producers may negatively affect both the models and the open web ecosystem in the long term. The Microsoft document cited in the dossier notes that it is unusual for a product to threaten the economics of its key content suppliers, and such a risk arises for large language models. In his statement this year, Microsoft CEO Satya Nadella is reported to have said that content behind the paywall should be licensed if it will be used for model training or supporting artificial intelligence systems with information. Nadella also stated that if he learned that OpenAI was collecting content behind the paywall without permission and using it to train models, Microsoft could raise its right to request that the models be retrained.

Internal correspondence in the file attributed to Nick Turley, OpenAI’s manager responsible for ChatGPT, also points to the economic risk facing publishers. It is stated that Turley said that chatbots can partially replace the content offered by publishers and that this effect may increase as the models develop. It is reported that OpenAI President Greg Brockman described the models as “very successful” in terms of news, and Nadella admitted that chatbots can reduce the need to go to the source website by allowing the user to access information directly through the artificial intelligence interface. The Times uses these statements to support its thesis that AI products not only transform existing content, but can directly compete with original content in certain use cases. Another Microsoft document states that there is a “real risk” that generative AI could seriously impact the employment of the people who produce the data on which the underlying models are trained.

One of the most striking headlines in the newly opened files is about the amount of data used. According to the court filing, the datasets used in the intermediate stages of OpenAI’s model training include copies of more than 91,692 works published by The New York Times, Daily News and the Center for Investigative Reporting. Another Common Crawl-based dataset is claimed to contain more than two million documents from the nytimes.com domain alone. The Times’ application also includes Brent Hecht’s description of data use on this scale in an internal memo dated January 2023 as “a staggering theft on an unprecedented scale” and “the largest labor theft in human history.” Since additional documents containing the original context of these words are not publicly available, it is not possible to see from the current file in which discussion context all the statements were used.

The filing also includes more detailed allegations about how OpenAI and Microsoft collect publisher content. According to the filing, OpenAI forwarded the entire GPT-3 training dataset to Microsoft, which used the data to evaluate how it could integrate OpenAI models into its own commercial products. It is claimed that Microsoft provided training data to OpenAI through projects called Project Taxi and Project Mango, and that there were copies of at least 160,903 unique works belonging to news organizations in the dataset created within the scope of Project Mango. The Times’ application also states that OpenAI employees discussed methods to bypass paywalls without being detected. It is claimed that OpenAI researcher Nick Ryder informed Brockman that he had found a method to bypass The New York Times paywall, and Brockman responded positively.

It is also claimed in the documents that OpenAI training sets such as WebText and WebText2 are highly based on scanned news content, and that millions of articles are collected from the Common Crawl archive. The Times also suggests that some researchers are working to remove copyright notices from training data on the grounds that they do not want the model to generate copyright notices for users. Steven Lieberman, an attorney pursuing the case for the New York Daily News, argues that the new materials show OpenAI and Microsoft knew their practices were problematic. However, a significant portion of the statements in question come from the legal filing filed by the plaintiffs, and since not all of the underlying documents are yet publicly available, the companies’ comprehensive context for these quotes cannot be seen. In this respect, the case directly concerns not only what content was used in model training in the past, but also within what legal limits the licensing, traffic and revenue sharing discussions between artificial intelligence companies and news publishers will be shaped.

Danish Kapoor