OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web
First reported by The Verge ·
You now know that OpenAI and Microsoft internally acknowledged their AI's potential to harm publishers and the web, yet proceeded with development.
Unsealed court documents from The New York Times' lawsuit against OpenAI and Microsoft reveal internal admissions that their AI models could harm the web and publishers. Microsoft's Director of Applied Science, Brent Hecht, characterized the scraping of web data for AI training as the "largest theft of labor in human history" and a "complete mockery of the idea of fair use." An internal Microsoft document stated that their AI content strategy initiated a "doom loop" that would damage both their models' performance and the web. OpenAI co-founder Greg Brockman expressed interest in the significant financial gains from commercial AI. Despite acknowledging that paywalled content should be licensed, OpenAI representatives were reportedly unaware of efforts to detect or remove such content from training data. Internal OpenAI discussions also noted ChatGPT's tendency to reproduce copyrighted material verbatim, admitting GPT-4 was "insanely good at regurgitation." Microsoft recognized that their extensive web scraping was not intended or compensated by content creators and admitted that LLMs can be a product that "destroys its own supply chain" by substituting for training data, potentially causing a 60 percent drop in search referral traffic.
These revelations signal a critical inflection point for content creators and the AI industry, suggesting a deliberate disregard for the economic foundations of online publishing. The "doom loop" described indicates a potential future where AI models degrade the very web content they rely on, creating a self-destructive cycle. This internal awareness, contrasted with public messaging, raises significant questions about corporate responsibility and the long-term viability of content-driven businesses in the age of generative AI.
The unsealed documents suggest that the pursuit of AI advancement and profit actively undermined the content ecosystem that fuels it. This has direct implications for how content is valued, licensed, and ultimately, created. The potential for AI to cannibalize referral traffic and substitute for human labor implies a fundamental restructuring of online information economies, necessitating new models for compensation and content utilization.
AI-written summary. May contain errors.