Static

Big AI's content problem: Take the work, keep the money

First reported by The Register ·

The signal ●○○○ Compiled by AI from The Register, the single source so far
Why you might care

If you use AI-generated code or content, you may be introducing unknown legal risks and intellectual property violations into your projects.

What happened

Big AI companies are facing scrutiny over their use of copyrighted material for training large language models (LLMs), with accusations of taking content without proper licensing or compensation. A lawsuit filed by The New York Times against OpenAI and Microsoft highlights this issue, with unsealed court documents revealing internal statements from Microsoft acknowledging the potentially "astonishing theft of unprecedented proportions" and that LLMs could "destroy its supply chain." OpenAI's internal discussions also indicated a lack of effort to exclude paywalled content from training datasets. Separately, a Ninth Circuit ruling in Doe v. GitHub offered a narrow win for GitHub, Microsoft, and OpenAI regarding copyright management information in AI-generated code, but did not resolve broader questions about fair use or license compliance for AI training data.

What it means

The core of the problem is that Big AI's business model appears to be predicated on absorbing vast amounts of existing work without adequately compensating creators, leading to a "doom loop" where AI development undermines its own content supply chain. Internal communications suggest awareness within companies like Microsoft that this practice could be detrimental, yet the pursuit of AI advancement continues unabated. This raises significant questions about the long-term sustainability of AI development if it relies on practices that deplete the very sources it needs to function and evolve.

The legal battles, such as the one initiated by The New York Times and the Doe v. GitHub case, are beginning to define the boundaries of AI's interaction with intellectual property. While rulings may offer temporary wins for AI companies on specific legal technicalities, they do not resolve the fundamental ethical and practical challenges of AI training data sourcing and creator compensation. The ongoing litigation and evolving legal landscape indicate a period of intense legal and contractual renegotiation for the entire AI industry, with potential ramifications for open-source licenses and intellectual property law.

AI-written summary. May contain errors.