Signal

NYT court filing: ChatGPT's head wrote that publishers face an "existential threat" and a Microsoft executive called AI training "an astonishing theft"

First reported by Ft ·

The signal ●●●○ Compiled by AI from Ft, Techmeme, The Information, Bloomberg Law and The Wrap
Why you might care

AI models will now require licensing fees for training data, increasing their development costs.

What happened

Internal communications from OpenAI and Microsoft leaders, recently unsealed in a New York Times lawsuit, reveal candid admissions about the threat generative AI poses to news publishers. Microsoft CEO Satya Nadella testified that AI chatbot conversations have "substituted" for publisher websites, and OpenAI's senior executive Nick Turley stated, "Our products are largely substitutive, period." During discovery, Microsoft's Director of Applied Science, Dr. Brent Hecht, described the unauthorized scraping of millions of news articles as potentially the "largest theft of labor in human history." OpenAI President Greg Brockman was also reportedly pleased with a method to circumvent the Times' paywall for scraping purposes. These statements emerged as part of a multidistrict litigation concerning the alleged illegal use of news articles for AI training, with the Trump administration arguing that such training constitutes fair use.

What it means

These internal statements, particularly the acknowledgment of "substitutive" products and "astonishing theft," directly contradict public arguments that AI models merely augment or transform existing content without directly competing with original publishers. This suggests a potential legal vulnerability for AI developers in copyright infringement cases, as these admissions could be used as evidence of intent and awareness of the harm caused to content creators. The unsealed documents imply that the core business models of news organizations are directly threatened by AI's ability to replicate and serve up information previously only accessible through subscriptions or direct site visits, potentially devaluing their content.

The admissions also signal a potential shift in how AI companies will need to approach data acquisition for training future models. If direct scraping and use of copyrighted material are deemed infringing, and if licensing becomes the norm as suggested by Nadella's testimony regarding paywalled content, the cost of developing and maintaining large language models could significantly increase. This could lead to greater consolidation in the AI market, favoring companies with the resources to negotiate extensive licensing deals, and potentially slowing down the pace of innovation or leading to more expensive AI products for consumers and businesses alike.

AI-written summary. May contain errors.