AI News AI资讯 2h ago Updated 1h ago 更新于 1小时前 49

Why is Anthropic destroying books? Anthropic为何在销毁书籍?

Anthropic's "Project Panama" was an internal initiative to destructively scan all the books in the world to build a high-quality training dataset for Claude, created before 2022 to avoid AI-contaminated text The company chose destructive scanning over obtaining copyright permissions or using pirated sources, arguing it fell under the "fair use" doctrine by transforming physical books into digital format and then discarding the originals A Northern California federal judge ruled that using propri Anthropic内部项目"Project Panama"涉及对全球书籍进行破坏性扫描,以获取高质量训练数据改进Claude模型 公司曾考虑使用盗版来源,后因法律风险转向破坏性扫描方案,并为此支付15亿美元和解金 法院裁定使用专有材料训练LLM本身不构成版权侵权,将"训练"等同于人类学习过程 扫描过程涉及拆解书籍、数字化后销毁原件,形成完整的物流供应链 文章质疑这种模式若被广泛采用,将对人类文化遗产保护构成威胁

72
Hot 热度
68
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • Anthropic's "Project Panama" was an internal initiative to destructively scan all the books in the world to build a high-quality training dataset for Claude, created before 2022 to avoid AI-contaminated text
  • The company chose destructive scanning over obtaining copyright permissions or using pirated sources, arguing it fell under the "fair use" doctrine by transforming physical books into digital format and then discarding the originals
  • A Northern California federal judge ruled that using proprietary material to train an LLM does not inherently constitute copyright infringement, equating AI training to human learning
  • Anthropic previously paid a $1.5 billion out-of-court settlement to authors, indicating the company had already faced legal consequences for using pirated books
  • The practice raises profound cultural and legal questions about the destruction of physical texts as a routine method for AI data procurement, with virtually no regulatory framework governing the treatment of books as cultural heritage

Why It Matters

This case represents a landmark legal and ethical moment for the AI industry, establishing a precedent that could normalize the physical destruction of cultural artifacts for data collection. For AI practitioners and researchers, it highlights the urgent need to develop sustainable, legally compliant data sourcing strategies before regulatory frameworks catch up. The ruling's equivalence of AI training to human learning could fundamentally reshape copyright law as it applies to generative AI.

Technical Details

  • Project Panama was Anthropic's codename for a large-scale destructive scanning operation aimed at procuring pre-2022 human-authored texts to train Claude, based on the premise that books contain "well-curated facts, well-organized analyses, and captivating fictional narratives" superior to post-2022 internet text contaminated by AI-generated content
  • The scanning process involved hiring specialized vendors who stripped books from bindings, cut pages to size, scanned them into digital form, and discarded the paper originals — a logistics operation managed by an experienced logistics manager with warehouse storage and labeled shelving
  • Anthropic initially attempted to use pirated sources before abandoning that approach "for legal reasons," then turned to destructive scanning as a workaround that leveraged the US "fair use" doctrine's transformative use provision
  • The court case Bartz v Anthropic PBC (Northern California District Court, late July 2025) resulted in a ruling that training an LLM on copyrighted material does not inherently constitute copyright infringement, with Judge William Alsup drawing an analogy between AI training and human education
  • Anthropic committed to creating a "forever" research library from the scanned materials, though the court noted no evidence the digital copies were shown, shared, or sold outside the company

Industry Insight

  • The legal precedent set in this case could trigger a wave of similar destructive scanning initiatives across the generative AI industry, creating an urgent need for companies to invest in licensed data pipelines and alternative data sourcing strategies before regulations restrict this practice
  • The $1.5 billion settlement and this ruling together signal that while current US copyright law favors AI developers, the legal landscape is unstable — companies should prepare for potential legislative changes that could criminalize or heavily regulate destructive data procurement methods
  • The cultural and ethical backlash against destroying physical books for AI training suggests that the industry should proactively engage with authors, publishers, and cultural institutions to establish cooperative data licensing models rather than exploiting legal loopholes that treat cultural heritage as an unregulated resource

TL;DR

  • Anthropic内部项目"Project Panama"涉及对全球书籍进行破坏性扫描,以获取高质量训练数据改进Claude模型
  • 公司曾考虑使用盗版来源,后因法律风险转向破坏性扫描方案,并为此支付15亿美元和解金
  • 法院裁定使用专有材料训练LLM本身不构成版权侵权,将"训练"等同于人类学习过程
  • 扫描过程涉及拆解书籍、数字化后销毁原件,形成完整的物流供应链
  • 文章质疑这种模式若被广泛采用,将对人类文化遗产保护构成威胁

为什么值得看

这篇文章揭示了AI训练数据获取背后的伦理与法律争议,对从业者理解版权边界和数据合规至关重要。同时,它提出了一个关键问题:当AI开始大规模消耗人类文化成果时,如何保护创作者权益和文化遗产。

技术解析

  • Anthropic采用"破坏性扫描"技术路线,通过拆解书籍、数字化后销毁原件的方式获取训练数据,规避版权许可流程
  • 项目涉及完整的物流体系:仓储管理、供应商协调、数字化处理等环节,形成工业化数据生产链
  • 法院将AI训练类比为人类学习过程,认为这属于合理使用范畴,不直接构成版权侵权
  • 数据筛选标准明确:优先选择2022年前创作的高质量文本,避免AI生成内容的污染
  • 扫描过程由专业供应商执行,包括拆书、裁切、扫描等工序,原件最终被销毁

行业启示

  • AI公司需要建立更透明的数据获取机制,避免依赖灰色地带的版权策略,否则将面临日益增长的法律诉讼风险
  • 行业应重新审视训练数据的伦理边界,平衡技术创新与创作者权益,探索授权合作的新模式
  • 文化遗产保护面临新挑战,需要建立针对AI时代的文化资产保护框架,防止人类知识成果被无偿消耗

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Claude Claude Training 训练 Dataset 数据集 Ethics 伦理 LLM 大模型