Microsoft exec says AI training might be the "largest labor theft in human history
First reported by Techspot ·
The core dispute over AI training data could lead to new licensing frameworks or significant legal judgments impacting AI development costs.
Internal documents reveal a Microsoft executive, Brent Hecht, described the training of large language models on internet data as potentially the "largest labor theft in human history." This statement emerged in court documents related to a copyright lawsuit filed by The New York Times against OpenAI and Microsoft. The tech companies publicly maintain that using copyrighted material for AI training constitutes fair use, akin to a student's reading. However, Hecht's private assessment suggests an acknowledgment of the controversial nature of their data acquisition practices. OpenAI executive Greg Brockman also admitted that ChatGPT can verbatim complete sentences from New York Times articles. OpenAI has previously stated that training AI without copyrighted material is impossible, and a former Meta executive noted that AI would cease to exist if copyright laws were strictly enforced against it.
This internal contradiction between public defense and private assessment highlights the escalating legal and ethical challenges faced by AI developers concerning intellectual property. The potential for AI models to replicate copyrighted content verbatim, as admitted by an OpenAI executive, directly undermines the 'fair use' defense and could force a reevaluation of how training data is sourced and compensated.
The news industry, in particular, faces an existential threat as AI-generated summaries can divert traffic from news websites, impacting revenue and the sustainability of journalism. The future of AI development may hinge on a resolution to these copyright battles, potentially leading to increased costs for model training and shifts in data acquisition strategies across the industry.
AI-written summary. May contain errors.