Static

Unsealed Briefs in Authors’ Case v. Microsoft/OpenAI

First reported by Authorsguild ·

The signal ●○○○ Compiled by AI from Authorsguild and Hacker News
Why you might care

Court documents show executives knew their AI training data was stolen and would put writers out of work.

What happened

New court filings in the Authors Guild v. OpenAI lawsuit reveal that executives at both OpenAI and Microsoft were aware that their use of copyrighted books for AI training was illegal and would negatively impact authors. Documents show OpenAI employees discussed the potential for their AI models to "make people unemployed" and "substitute for the labor of people." One OpenAI researcher, Tarun Gogineni, even stated his mission was to have GPT models write the "last two books of [George R.R. Martin's] A Song of Ice and Fire" series, viewing authors' concerns about stolen datasets and job losses as "acceptable economic disruption." Microsoft was reportedly aware of OpenAI's use of sources like LibGen, a website known for hosting pirated books, as early as April 2019. In an effort to conceal this activity, OpenAI launched a "Project Clear" initiative in the summer of 2022 to delete LibGen files from its systems. The filings highlight a deliberate disregard for copyright law and the potential consequences for the writing profession, with internal discussions prioritizing the race for AI advancement over legal and ethical considerations.

What it means

These unsealed documents signal a critical shift in the narrative of AI development, moving beyond technical advancements to expose the ethical and legal compromises made by leading companies. The explicit acknowledgment of potential job displacement and the intentional use of pirated material suggests a calculated business strategy that prioritized market dominance over creator rights. This could lead to increased scrutiny from regulators and a demand for greater transparency in AI training data sources across the industry. The revelations also have significant implications for the future of creative industries. If AI models are built on a foundation of intellectual property theft, it undermines the value of original human creation and could disincentivize authors from producing new works. The Authors Guild's legal action, now centered on demonstrable internal knowledge of wrongdoing, could set a precedent for how copyright infringement in AI training is addressed, potentially leading to demands for compensation and changes in how AI companies acquire and use data.

The explicit admission of potential job losses and the deliberate choice to ignore these concerns indicate a business model that may be fundamentally at odds with supporting creative professions. This puts pressure on existing legal frameworks and raises questions about the sustainability of an AI industry built on potentially exploitative practices. Future AI development may need to incorporate ethical data acquisition and demonstrate a clearer path to coexisting with, rather than replacing, human creators. The focus on OpenAI and Microsoft's internal knowledge of their actions is a pivotal moment, potentially shaping how future AI lawsuits are framed and prosecuted. It suggests that the "move fast and break things" mentality, when applied to intellectual property, carries significant legal and reputational risks. The industry will likely watch closely to see if this leads to stricter industry self-regulation or more assertive legislative and judicial intervention regarding AI training data.

whyItMattersOnceMoreJustInCaseThisIsTheFirstThingTheyReadFromThisOutputAtTheTopOfTheCardIImagineThemSittingThereLookingAtTheCard:

AI-written summary. May contain errors.