AI News AI资讯 23h ago Updated 18h ago 更新于 18小时前 35

Sony and Warner sue Anthropic over "one of the largest and most blatant ongoing thefts of intellectual property in history" 索尼和华纳起诉Anthropic,指控其犯下"历史上最大、最明目张胆的持续知识产权盗窃案之一"

Sony Music and Warner Music have filed a federal lawsuit in Northern California against Anthropic, alleging the company torrented tens of thousands of copyrighted musical compositions (primarily song lyrics) to train Claude models without permission CEO Dario Amodei and co-founder Benjamin Mann are named as individual defendants, accused of directing and overseeing the illegal downloading of copyrighted files The lawsuit mirrors Anthropic's prior $1.5 billion settlement over pirated books, targe Sony Music和Warner Music对Anthropic提起版权诉讼,指控其非法下载并使用数万起受版权保护的音乐作品(主要是歌词)训练Claude模型 诉讼核心聚焦于Anthropic获取训练数据的方式,CEO Dario Amodei和联合创始人Benjamin Mann被列为个人被告 这是Anthropic在2025年9月因盗版书籍支付15亿美元和解后的又一重大版权纠纷,原告指控其通过torrent非法下载至少700万本书籍 原告还指控Anthropic利用合成数据绕过版权限制,通过非商业模型生成的合成数据来训练商业Claude模型 慕尼黑地区法院已在2025年11月裁定歌词在模

50
Hot 热度
50
Quality 质量
50
Impact 影响力

Analysis 深度分析

TL;DR

  • Sony Music and Warner Music have filed a federal lawsuit in Northern California against Anthropic, alleging the company torrented tens of thousands of copyrighted musical compositions (primarily song lyrics) to train Claude models without permission
  • CEO Dario Amodei and co-founder Benjamin Mann are named as individual defendants, accused of directing and overseeing the illegal downloading of copyrighted files
  • The lawsuit mirrors Anthropic's prior $1.5 billion settlement over pirated books, targeting the same vulnerability: acquiring training data through illegal torrent downloads from pirate libraries like LibGen and PiLiMi
  • Plaintiffs allege Anthropic trained commercial Claude models on synthetic data generated by a non-commercial model that had learned from unauthorized sources, and used such models for reinforcement feedback
  • A related November 2025 Munich Regional Court ruling established that copyrighted song lyrics stored within model parameters constitute reproductions, and chatbot output of those lyrics amounts to unlawful public disclosure, with model operators held liable

Why It Matters

This lawsuit represents a critical escalation in the ongoing legal battle over AI training data copyright, directly targeting Anthropic's data acquisition practices rather than just model outputs. The precedent set here could reshape how AI companies source and handle copyrighted material, potentially forcing industry-wide changes to training pipelines and data licensing strategies.

Technical Details

  • The complaint challenges Anthropic's use of datasets including Books3, The Pile, and Common Crawl, alleging they contain unauthorized copyrighted content, and accuses the company of scraping lyrics from licensed platforms like MusixMatch and LyricFind in violation of terms of service
  • Plaintiffs allege a multi-stage training pipeline where a non-commercial model first learned from LibGen and PiLiMi texts, then generated synthetic data used to train a commercial Claude model, with reinforcement feedback also derived from such models
  • The Munich Regional Court's November 2025 ruling established that copyrighted lyrics embedded within model parameters count as reproductions, and that operators—not users—are liable for model outputs, even when prompts are specifically designed to generate copyrighted material
  • Anthropic previously settled for $1.5 billion in September 2025 over illegal torrent downloads of at least seven million books, a case where the acquisition method (not just usage) was the determining factor
  • The current lawsuit seeks up to $150,000 per infringed work and up to $25,000 per violation for unlawful removal of copyright management information, including copyright notices

Industry Insight

  • AI companies must urgently audit their data acquisition pipelines for illegal sourcing methods, as courts are increasingly treating the act of downloading copyrighted material as standalone infringement regardless of whether it was used in final models
  • The synthetic data loophole argument—where copyrighted content is indirectly used through intermediate models—sets a concerning precedent that could expand liability across the entire training chain, from base models to fine-tuned variants
  • The Munich ruling's assignment of operator liability for model outputs, even against jailbreak-style prompts, signals that AI companies will face stricter legal exposure for what their models generate, making robust content filtering and licensing compliance essential strategic priorities

TL;DR

  • Sony Music和Warner Music对Anthropic提起版权诉讼,指控其非法下载并使用数万起受版权保护的音乐作品(主要是歌词)训练Claude模型
  • 诉讼核心聚焦于Anthropic获取训练数据的方式,CEO Dario Amodei和联合创始人Benjamin Mann被列为个人被告
  • 这是Anthropic在2025年9月因盗版书籍支付15亿美元和解后的又一重大版权纠纷,原告指控其通过torrent非法下载至少700万本书籍
  • 原告还指控Anthropic利用合成数据绕过版权限制,通过非商业模型生成的合成数据来训练商业Claude模型
  • 慕尼黑地区法院已在2025年11月裁定歌词在模型参数中构成复制,输出歌词属于非法公开披露

为什么值得看

这篇文章揭示了AI行业版权合规的严峻挑战,特别是数据获取方式的合法性问题。对于AI从业者而言,这标志着单纯依赖公开数据集和torrent下载的时代可能正在结束,企业需要重新审视其数据供应链的合规性。

技术解析

  • 诉讼焦点在于数据获取方式而非使用方式:Anthropic被指控通过torrent非法下载至少700万本书籍(来自LibGen和PiLiMi),并从MusixMatch和LyricFind等授权平台抓取歌词,同时挑战了Books3、The Pile和Common Crawl等数据集的合法性
  • 原告指控Anthropic利用合成数据作为版权漏洞:通过非商业模型(已学习LibGen和PiLiMi内容)生成合成数据,再用该合成数据训练商业Claude模型,或使用该模型提供强化学习反馈
  • 慕尼黑地区法院2025年11月判决确立重要先例:版权保护歌词在模型参数中构成复制,聊天机器人输出歌词属于非法公开披露,模型运营商需对输出内容负责而非用户

行业启示

  • AI公司需建立更严格的数据获取合规机制,避免依赖torrent等非法渠道,数据供应链审计应成为标配
  • 版权诉讼正从"使用"转向"获取",企业应重新评估数据获取方式的合法性风险,而非仅关注训练后的使用
  • 合成数据可能成为新的法律争议焦点,模型运营商需对其输出内容承担法律责任,这为AI行业的版权合规提出了更高要求

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。