AI News AI资讯 6h ago Updated 3h ago 更新于 3小时前 49

Microsoft says virtually nobody was grabbing NYT articles through its chatbot 微软称几乎无人通过其聊天机器人抓取纽约时报文章

Microsoft disclosed 8.2 million Copilot chat logs in copyright litigation, claiming only 59,545 contained 16+ matching words with news content and just 24 responses matched 30+ words with book content The New York Times strongly disputes Microsoft's findings, asserting that Copilot and OpenAI's systems directly substitute for journalistic work and undermine the industry Microsoft argues the low reproduction rate supports its fair use defense, contending that LLM training is a transformative use Microsoft在版权诉讼中提交820万条Copilot聊天记录,显示极少复制新闻文章或书籍原文 分析显示59,545条记录包含至少16个匹配词,212本书中仅10本有匹配,作者诉讼专家仅发现24条含30个以上匹配词 Microsoft主张AI训练使用版权内容属于"合理使用",系统用途与原作显著不同且具有变革性 The New York Times等原告强烈反对,指控Microsoft和OpenAI盗用内容并直接竞争 案件正寻求Summary Judgment早期结案,Trump政府提交利益声明支持OpenAI

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Microsoft disclosed 8.2 million Copilot chat logs in copyright litigation, claiming only 59,545 contained 16+ matching words with news content and just 24 responses matched 30+ words with book content
  • The New York Times strongly disputes Microsoft's findings, asserting that Copilot and OpenAI's systems directly substitute for journalistic work and undermine the industry
  • Microsoft argues the low reproduction rate supports its fair use defense, contending that LLM training is a transformative use regardless of occasional text overlap
  • The Trump administration filed a statement of interest supporting OpenAI in the NYT case, adding political weight to Microsoft's position
  • The consolidated lawsuit seeks summary judgment from Microsoft, which could end the case early if the judge agrees with the defendant

Why It Matters

This case represents one of the most significant copyright battles in AI history, directly shaping whether companies can legally train large language models on copyrighted material without permission or compensation. The outcome will establish precedents affecting the entire AI industry's training practices and the economic viability of news publishers and authors whose works power these systems.

Technical Details

  • Microsoft analyzed 8.2 million Copilot chat logs selected for keywords related to News Plaintiffs' websites, finding 59,545 instances with at least 16 words matching grounded news content
  • An expert for the Center for Investigative Reporting identified only 51 instances of "substantial overlap" with CIR work across the entire dataset
  • In the authors' suit, only 24 of 8.2 million conversations contained 30+ matching words, and merely 10 of 212 evaluated books showed any matches at all
  • Copilot's architecture relies on grounding responses using copyrighted training data, but Microsoft maintains the output serves a fundamentally different purpose than the original works

Industry Insight

  • The extremely low reproduction rate (roughly 0.7% of logs showing any overlap) suggests current LLMs do not function as direct content substitutes, which could strengthen fair use arguments but may not satisfy plaintiffs seeking structural remedies
  • Government intervention supporting OpenAI signals potential executive branch alignment with the AI industry, which could influence judicial outcomes and set a policy tone for future cases
  • Publishers and authors should consider that litigation alone may not prevent model training; legislative or licensing-based solutions may offer more durable protection for creative works in the AI era

TL;DR

  • Microsoft在版权诉讼中提交820万条Copilot聊天记录,显示极少复制新闻文章或书籍原文
  • 分析显示59,545条记录包含至少16个匹配词,212本书中仅10本有匹配,作者诉讼专家仅发现24条含30个以上匹配词
  • Microsoft主张AI训练使用版权内容属于"合理使用",系统用途与原作显著不同且具有变革性
  • The New York Times等原告强烈反对,指控Microsoft和OpenAI盗用内容并直接竞争
  • 案件正寻求Summary Judgment早期结案,Trump政府提交利益声明支持OpenAI

为什么值得看

这篇文章揭示了AI大模型版权争议的核心法律博弈,Microsoft通过大规模数据展示其系统极少直接复制受版权保护内容,试图论证训练使用的"合理使用"性质。这对整个AI行业具有标志性意义,将直接影响大模型训练数据的法律边界、合规成本和商业模式可持续性。

技术解析

  • Microsoft向出版商聘请的专家提供了820万条Copilot聊天记录,这些数据被特意筛选为最可能包含原告网站关键词的内容
  • 分析结果显示:59,545条记录包含至少16个与新闻内容匹配的词;CIR专家发现51处"实质性重叠";作者诉讼专家发现24条包含至少30个匹配词
  • 在212本评估的书籍中,只有10本有任何匹配,表明书籍内容的复制率极低
  • Microsoft主张AI系统虽然依赖版权材料,但用途与原作显著不同,"几乎不削弱LLM训练的变革性目的"

行业启示

  • 版权争议将成为AI行业持续面临的法律风险,企业需要建立更完善的内容筛选和版权合规机制,避免直接复制受保护内容
  • "合理使用"原则在AI训练场景下的适用性将逐步通过司法实践明确,可能重塑行业数据使用规范和训练方法论
  • 大型科技公司正通过法律策略(如寻求Summary Judgment)加速解决版权争议,行业需密切关注判决结果对商业模式和竞争格局的深远影响

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Closed Source 闭源 LLM 大模型 Legal AI 法律AI Regulation 监管 Policy 政策