Microsoft Chat Logs Reveal Minimal The New York Times Article Retrieval in Copilot Discovery
In legal filings submitted on Friday, Microsoft reported that an analysis of 8.2 million Copilot chat logs showed virtually no systemic reproduction of articles from The New York Times. As part of ongoing copyright litigation consolidated in federal court, Microsoft argues that large language model training constitutes protected fair use.
The Bottom Line
- Discovery Dataset: Microsoft reviewed 8.2 million Copilot chat logs specifically targeted by keyword filters to unearth potential matches with news publishers’ sites.
- Limited Overlap: Only 59,545 conversations contained at least 16 words matching training content from news providers, with similar sparse results found in evaluations of book authors’ works.
- Legal Stakes: The findings form the basis of Microsoft’s motion for summary judgment, while plaintiffs maintain that the technology relies on misappropriated journalism.
Parsing the Discovery Data in the Copyright Litigation
The core dispute centers around how generative artificial intelligence models process and reproduce copyrighted material. According to legal filings, Microsoft disclosed the analysis of 8.2 million Copilot chat logs. The company selected these specific logs because they hit keywords implicating the use of news plaintiffs’ websites, representing the highest probability segment for finding overlapping text.
Out of those millions of interactions, Microsoft’s analysis indicated that 59,545 conversations contained at least 16 words in common with news content utilized to ground the artificial intelligence model. In a parallel assessment involving book authors, experts identified only 24 Copilot responses containing at least 30 matching words across the entire dataset. Furthermore, out of 212 evaluated books, only 10 contained any matches at all. An expert working for the Center for Investigative Reporting also identified 51 instances of substantial overlap within the dataset.
These figures form the quantitative foundation of Microsoft’s defense. The tech giant contends that these minor instances of text reproduction do not substitute for the original journalism. Therefore, the company argues, training large language models on copyrighted works qualifies as a transformative use under copyright law.
The Plaintiffs’ Counter-Arguments and Regulatory Interventions
Legal counsel for the publishing houses firmly rejected Microsoft’s characterization of the discovery data. Ian Crosby, lead counsel for The New York Times, stated in a statement that the evidence demonstrates commercial products were built by taking journalism without authorization, which directly threatens the publishing business model. Copyright claims have also been brought forward by organizations like Newsday and the Seattle Times against both Microsoft and its partner OpenAI.
The Trump administration recently filed a statement of interest supporting OpenAI within the New York Times litigation.
Comparative Metrics of Copilot Discovery Findings
| Dataset Evaluated | Total Volume Examined | Identified Matches / Overlaps |
|---|---|---|
| News Publisher Keyword Logs | 8,200,000 conversations | 59,545 chats (≥ 16 matching words) |
| Center for Investigative Reporting | Subset of 8.2M logs | 51 instances of substantial overlap |
| Book Authors Evaluation | 8,200,000 conversations | 24 responses (≥ 30 matching words across 212 books) |
Market Implications for Enterprise Artificial Intelligence
Disclaimer: The information provided in this article is for educational and informational purposes only and does not constitute financial advice.
