AI News AI资讯 18h ago Updated 15h ago 更新于 15小时前 62

US government sides with OpenAI on issue of training LLMs on copyrighted material 美国政府支持OpenAI,在LLM使用版权材料训练问题上

The Trump administration filed a 20-page amicus brief defending OpenAI's unlicensed use of copyrighted material to train LLMs, citing the need to maintain U.S. global AI leadership The brief argues that constraining LLM development under a narrow interpretation of fair use would hinder creative and scientific progress and American economic prosperity The core legal debate centers on whether AI training constitutes "transformative" fair use, with comparisons drawn between how LLMs process works v 特朗普政府提交20页法庭之友简报,明确支持OpenAI未经许可使用版权材料训练LLM,主张此举关乎美国AI全球领导地位 案件核心争议聚焦合理使用原则,关键判断标准在于AI训练是否构成“变革性”使用而非简单复制 先例显示法院倾向保护AI训练行为本身(如Anthropic案),但数据来源非法性(如影子图书馆)仍会导致高额赔偿 政府简报虽无直接管辖权,但反映行政层面对AI产业发展的政策倾斜,可能影响诉讼走向与行业预期 该诉讼结果将确立AI行业使用版权数据的合法性边界,对模型训练策略与数据合规产生深远影响

75
Hot 热度
70
Quality 质量
72
Impact 影响力

Analysis 深度分析

TL;DR

  • The Trump administration filed a 20-page amicus brief defending OpenAI's unlicensed use of copyrighted material to train LLMs, citing the need to maintain U.S. global AI leadership
  • The brief argues that constraining LLM development under a narrow interpretation of fair use would hinder creative and scientific progress and American economic prosperity
  • The core legal debate centers on whether AI training constitutes "transformative" fair use, with comparisons drawn between how LLMs process works versus human readers
  • Previous rulings have largely favored AI companies; Judge William Alsup previously noted Anthropic's training was transformative, with the $1.5 billion settlement stemming from piracy issues rather than copyright infringement itself
  • The brief carries strategic weight despite lacking jurisdiction in the Southern District of New York case

Why It Matters

The Trump administration's intervention signals high-level governmental support for the AI industry's training practices, potentially shaping the legal landscape for all major AI companies operating in the U.S. This could establish a precedent that either protects or constrains how the industry accesses copyrighted data, directly impacting training strategies and costs.

Technical Details

  • LLMs (ChatGPT, Claude, Gemini) are trained on massive databases of copyrighted published works including books, articles, and media, scraped without permission from rights holders
  • The fair use doctrine is the central legal framework, specifically whether AI training qualifies as "transformative" use under copyright law
  • Judge Alsup's prior ruling compared LLM training to a human reader studying works to become a writer, emphasizing creation of something different rather than replication
  • The Anthropic case established a distinction between copyright infringement (training on copyrighted material) and illegal sourcing (using shadow libraries), with the latter triggering the $1.5 billion settlement
  • The NYT v. OpenAI case is being tried in U.S. District Court for the Southern District of New York, where the Trump administration's brief serves as an amicus curiae submission

Industry Insight

  • AI companies should anticipate continued governmental backing for their training practices, but publishers may pursue alternative legal strategies beyond fair use arguments
  • The distinction between lawful training and unlawful data sourcing (as seen in the Anthropic case) means companies must audit their data pipelines for compliance, not just fair use claims
  • The regulatory environment is increasingly politicized; companies should engage with policy advocates and monitor executive orders that could shift the legal framework for AI training data access

TL;DR

  • 特朗普政府提交20页法庭之友简报,明确支持OpenAI未经许可使用版权材料训练LLM,主张此举关乎美国AI全球领导地位
  • 案件核心争议聚焦合理使用原则,关键判断标准在于AI训练是否构成“变革性”使用而非简单复制
  • 先例显示法院倾向保护AI训练行为本身(如Anthropic案),但数据来源非法性(如影子图书馆)仍会导致高额赔偿
  • 政府简报虽无直接管辖权,但反映行政层面对AI产业发展的政策倾斜,可能影响诉讼走向与行业预期
  • 该诉讼结果将确立AI行业使用版权数据的合法性边界,对模型训练策略与数据合规产生深远影响

为什么值得看

本文揭示了美国政府对AI版权争议的政策立场,直接影响AI公司的训练数据获取策略与合规成本。法律先例的走向将决定整个行业能否继续使用现有版权材料进行模型训练,关乎技术创新与版权保护的平衡。

技术解析

  • LLM训练依赖规模庞大的版权材料数据库,涵盖书籍、文章及其他媒体内容,AI企业通常未经许可直接抓取使用这些数据。
  • 合理使用原则成为法律争议焦点,法院需评估AI训练是否具备“变革性”——即是否将原始作品转化为新功能或表达,而非简单替代原市场。
  • Anthropic案判决区分了训练行为与数据来源:法官认定LLM学习作品类似人类阅读后创作,不构成侵权;但因使用非法影子图书馆获取数据,仍需支付15亿美元赔偿。
  • 政府简报援引总统行政命令,强调AI发展对国家经济繁荣、科学进步及全球标准制定的战略价值,为产业辩护提供政策依据。

行业启示

  • AI公司应优先布局授权数据集、开源材料或合成数据,降低版权诉讼风险,并建立训练数据的合规审查机制。
  • 政策层面可能推动更清晰的AI训练版权框架,行业需积极参与立法讨论,塑造兼顾创新与版权保护的法律环境。
  • 类似诉讼预计将持续增多,企业应关注法院对“变革性使用”标准的细化,及时调整数据策略与商业模式。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

OpenAI OpenAI LLM 大模型 Training 训练 Policy 政策 Regulation 监管