Microsoft says virtually nobody was grabbing NYT articles through its chatbot
Microsoft disclosed 8.2 million Copilot chat logs in copyright litigation, claiming only 59,545 contained 16+ matching words with news content and just 24 responses matched 30+ words with book content The New York Times strongly disputes Microsoft's findings, asserting that Copilot and OpenAI's systems directly substitute for journalistic work and undermine the industry Microsoft argues the low reproduction rate supports its fair use defense, contending that LLM training is a transformative use
Analysis
TL;DR
- Microsoft disclosed 8.2 million Copilot chat logs in copyright litigation, claiming only 59,545 contained 16+ matching words with news content and just 24 responses matched 30+ words with book content
- The New York Times strongly disputes Microsoft's findings, asserting that Copilot and OpenAI's systems directly substitute for journalistic work and undermine the industry
- Microsoft argues the low reproduction rate supports its fair use defense, contending that LLM training is a transformative use regardless of occasional text overlap
- The Trump administration filed a statement of interest supporting OpenAI in the NYT case, adding political weight to Microsoft's position
- The consolidated lawsuit seeks summary judgment from Microsoft, which could end the case early if the judge agrees with the defendant
Why It Matters
This case represents one of the most significant copyright battles in AI history, directly shaping whether companies can legally train large language models on copyrighted material without permission or compensation. The outcome will establish precedents affecting the entire AI industry's training practices and the economic viability of news publishers and authors whose works power these systems.
Technical Details
- Microsoft analyzed 8.2 million Copilot chat logs selected for keywords related to News Plaintiffs' websites, finding 59,545 instances with at least 16 words matching grounded news content
- An expert for the Center for Investigative Reporting identified only 51 instances of "substantial overlap" with CIR work across the entire dataset
- In the authors' suit, only 24 of 8.2 million conversations contained 30+ matching words, and merely 10 of 212 evaluated books showed any matches at all
- Copilot's architecture relies on grounding responses using copyrighted training data, but Microsoft maintains the output serves a fundamentally different purpose than the original works
Industry Insight
- The extremely low reproduction rate (roughly 0.7% of logs showing any overlap) suggests current LLMs do not function as direct content substitutes, which could strengthen fair use arguments but may not satisfy plaintiffs seeking structural remedies
- Government intervention supporting OpenAI signals potential executive branch alignment with the AI industry, which could influence judicial outcomes and set a policy tone for future cases
- Publishers and authors should consider that litigation alone may not prevent model training; legislative or licensing-based solutions may offer more durable protection for creative works in the AI era
Disclaimer: The above content is generated by AI and is for reference only.