Microsoft court filings: an expert hired by publishers found that only ~60K of 8.2M Copilot chat logs contained at least 16 words in common with news content

Microsoft has submitted court filings in its copyright infringement lawsuit with publishers and authors, presenting data from an expert analysis of 8.2 million Copilot chat logs. The analysis, commissioned by Microsoft, reportedly found that only a small fraction of these logs (approximately 0.7%) contained at least 16 words in common with news content used to train the AI model. Microsoft argues these findings support its fair use defense, asserting that while copyrighted material is used for training, the AI's output is transformative and does not directly substitute for the original works. Publishers, including The New York Times, strongly dispute these conclusions, accusing Microsoft and OpenAI of "theft" and creating products that compete with and threaten their journalism. The filings are part of Microsoft's push for a summary judgment to potentially end the case early, while publishers and authors seek to continue the legal battle.

AI Signal Decode

Microsoft's legal defense hinges on a statistical analysis of Copilot's output, aiming to demonstrate minimal direct reproduction of copyrighted text. By highlighting that only a tiny percentage of 8.2 million sampled chat logs shared a significant word overlap (16+ words) with news content, Microsoft attempts to prove that its AI does not wholesale copy or substitute for journalistic or literary works. This data is presented to bolster the 'fair use' argument, which posits that the use of copyrighted material for training large language models is transformative and serves a different purpose than the original creation. The company asserts that occasional text reproduction does not negate this transformative intent.

The implications for the publishing industry are substantial, as this case could set a precedent for how AI models are trained and how copyright law applies in the digital age. Publishers fear that AI companies are profiting from their content without proper compensation or licensing, directly undermining their business models by potentially replacing the need for original reporting. Microsoft's data, if accepted by the court, could validate the AI industry's current training practices. Conversely, if publishers prevail, it could force significant changes in AI development, potentially requiring new licensing frameworks for training data and impacting the economic viability of AI development.

From a technical standpoint, the dispute centers on the definition of 'substantial overlap' and 'transformative use' in the context of generative AI. Critics argue that even small percentages of reproduced text can be damaging if that text is crucial or directly competitive. Microsoft's filings, while presenting specific word-count thresholds, do not detail the nature or context of the overlapping content. Future developments to watch include how the judge interprets the statistical evidence, the admission of further expert testimony, and whether the court will establish clearer guidelines for fair use in AI training, potentially influencing the development and deployment of future AI models across various creative industries.