The Seattle Times and Newsday sue OpenAI and Microsoft, alleging the companies trained AI on their journalism; Microsoft and OpenAI are funders of Seattle Times

The Seattle Times and Newsday have filed a lawsuit against Microsoft and OpenAI, alleging that their journalism was used without permission to train artificial intelligence models. The newspapers claim hundreds of thousands of articles were scraped, bypassing paywalls and ignoring terms of service, to develop AI products like ChatGPT and Copilot. This legal action seeks financial damages and the destruction of AI models trained on their copyrighted content. The lawsuit is particularly notable as Microsoft has previously funded certain Seattle Times journalism projects and jointly funded an AI fellowship with OpenAI that included both news organizations. Both plaintiffs and defendants have recently filed for summary judgment in a related New York Times case. This litigation highlights the ongoing tension between AI developers and content creators regarding the use of copyrighted material for AI training, with potential implications for the future of independent journalism and the AI industry.

AI Signal Decode

The core of the lawsuit revolves around allegations of copyright infringement, with The Seattle Times and Newsday contending that Microsoft and OpenAI illicitly obtained and utilized their published articles to train generative AI models. The suit specifically claims that the companies bypassed paywalls and disregarded terms of service, effectively scraping vast amounts of journalistic work without consent or compensation. This alleged practice threatens the financial viability of news organizations, as the generated AI content can directly compete with original reporting, potentially undermining the business models that support expensive, in-depth journalism. The plaintiffs are seeking both monetary damages and the removal of their content from existing AI training datasets and models.

The market implications of this lawsuit are significant, contributing to a broader legal battle over intellectual property in the age of AI. While OpenAI has entered into licensing agreements with several publishers, the success of this suit could pressure the company and others to pursue more comprehensive licensing strategies or face similar legal challenges. The involvement of Microsoft as both a funder and a defendant adds a layer of complexity, especially given the Seattle Times's acknowledgment of editorial independence despite prior funding. The outcome could set crucial precedents for how AI companies access and utilize copyrighted data, influencing investment in AI development and the sustainability of the news media industry.

From a technical perspective, the lawsuit underscores the challenges in identifying and attributing AI-generated content derived from specific training data. The Seattle Times provided an example of ChatGPT reproducing their reporting almost verbatim, highlighting the direct lineage some AI outputs can have to source material. The news organizations are asking for the destruction of AI models, a technically challenging request given the distributed nature of AI development and the difficulty in isolating and removing specific datasets retroactively. Future developments to watch include the legal rulings on summary judgment motions, the potential for a broader industry-wide settlement, and whether AI developers will implement more robust methods for tracking and compensating original content sources.