Microsoft exec called AI scraping 'the largest theft of labor in human history'
First reported by TechCrunch ·
AI models trained on copyrighted material may be retrained, impacting their capabilities and the data they were built upon.
New unredacted court filings reveal that a senior Microsoft executive described the AI training practices of companies like OpenAI as "the largest theft of labor in human history." These filings, part of a lawsuit by The New York Times against AI firms for copyright infringement, detail allegations of bypassing paywalls and mass scraping of content to build training datasets. Notably, these datasets allegedly contained over 91,000 works from the NYT and over 2 million documents from nytimes.com. Internal documents suggest that Microsoft's own AI products, like Copilot, have significantly reduced click-through rates to The New York Times' website, potentially harming the publishers' economic models. Microsoft CEO Satya Nadella also testified that paywalled content should be licensed for AI training, and he would have required OpenAI to retrain models if he knew paywalled content was used.
The internal admissions and Microsoft executive's statements indicate a potential shift in how AI companies approach data acquisition and fair use defense. The acknowledgment of "existential threat" and "doom loop" effects on publishers suggests that the industry may face increased scrutiny and regulatory pressure regarding the economic impact of AI on content creators. This could lead to new licensing frameworks or restrictions on data scraping, forcing AI developers to find alternative, potentially more expensive, training data sources.
The unsealed documents also highlight how AI developers deliberately stripped copyright notices, suggesting a clear awareness of the contentious nature of their data acquisition methods. The scale of alleged copyright infringement, with millions of articles and specific works included in training datasets, represents a significant challenge to the current "fair use" arguments. Future legal battles and policy discussions will likely focus on these specific actions and their implications for the established norms of intellectual property and digital content.
AI-written summary. May contain errors.