Static

Microsoft exec called AI scraping ‘the largest theft of labor in human history,’ new unredacted filings reveal

First reported by TechCrunch ·

The signal ●●○○ Compiled by AI from TechCrunch, the single source so far
Why you might care

AI training datasets will need to be licensed, increasing costs for AI developers and potentially limiting the scale of future models.

What happened

New unredacted court filings in The New York Times' lawsuit against OpenAI and Microsoft reveal internal communications where Microsoft executives described AI training practices as "the largest theft of labor in human history." The filings allege that OpenAI and Microsoft bypassed paywalls to scrape vast amounts of copyrighted content, including over 91,692 New York Times articles in OpenAI's mid-training datasets alone. Microsoft's internal documents show its Copilot "answer engine" reduced New York Times click-through rates by up to 93%, a decline described as a "doom loop" that would harm both AI models and the web. Microsoft CEO Satya Nadella testified that paywalled content should be licensed for AI training, and he would have required OpenAI to retrain models if aware of such scraping. OpenAI leadership acknowledged their models pose an "existential threat" to publishers, being "largely substitutive" of original work.

What it means

The internal admissions from Microsoft and OpenAI directly challenge the "fair use" defense by acknowledging the substitutive nature of AI outputs and the market harm inflicted on publishers. This suggests a potential shift in legal interpretations, moving away from broad claims of fair use for training data towards a requirement for explicit licensing, especially for content behind paywalls. The scale of data ingestion detailed, including bypassing paywalls and stripping copyright notices, indicates aggressive tactics that may face increased regulatory scrutiny.

This case directly affects the economic model for content creators and AI developers. If AI companies are forced to license all training data, the cost of developing large language models could significantly increase, potentially consolidating the market among well-funded entities. Publishers, on the other hand, may see a new revenue stream emerge but also face the challenge of how to effectively manage licensing for such massive datasets.

AI-written summary. May contain errors.