AI News AI资讯 3h ago Updated 2h ago 更新于 2小时前 50

AI companies are reportedly shredding books after using them to train AI models 据报道,人工智能公司在使用书籍训练AI模型后将其粉碎

AI companies are purchasing millions of secondhand books through intermediaries to source high-quality training data for their models, avoiding public backlash. Leading AI companies are turning to human-authored print sources that predate 2022 to avoid contamination from "AI slop." ISBNdb has pivoted its business to offer specialized services for bulk-purchasing books for AI companies, with orders ranging from 1,000 to one million books in a single transaction. Anthropic reportedly invested mill AI公司通过中间商大量采购二手书作为训练数据,以避开公众对网络“AI垃圾内容”的争议。 行业巨头如Anthropic和Google被指控利用数百万本版权书籍训练模型,部分书籍在数字化后被销毁。 ISBNdb等平台转向为AI企业提供批量购书服务,单次订单可达百万册,引发伦理与版权争议。 扫描过程多采用破坏性方式(如拆书),导致实体书永久消失,可能影响文化遗产保存。 尽管法律上“合理使用”原则支持AI训练,但大规模私有化知识资源引发对未来信息可及性的担忧。

75
Hot 热度
68
Quality 质量
72
Impact 影响力

Analysis 深度分析

TL;DR

  • AI companies are purchasing millions of secondhand books through intermediaries to source high-quality training data for their models, avoiding public backlash.
  • Leading AI companies are turning to human-authored print sources that predate 2022 to avoid contamination from "AI slop."
  • ISBNdb has pivoted its business to offer specialized services for bulk-purchasing books for AI companies, with orders ranging from 1,000 to one million books in a single transaction.
  • Anthropic reportedly invested millions in extracting information from printed books to build its Claude AI models and then destroyed them, leading to a $1.5 billion fine for maintaining a repository of pirated books.
  • The destruction of books raises ethical concerns about the removal of rare or out-of-print titles from circulation and the lack of public access to scanned content used for AI training.

Why It Matters

This trend highlights the growing demand for high-quality, uncontaminated data in AI development, which is crucial for advancing model performance and reliability. However, it also raises significant ethical and legal questions regarding copyright, the preservation of literary history, and the potential loss of valuable cultural artifacts. Understanding these dynamics is essential for AI practitioners, researchers, and policymakers to navigate the complex landscape of data sourcing and intellectual property rights.

Technical Details

  • Data Quality: AI companies prioritize high-quality, human-authored content over AI-generated content to avoid contamination and improve model accuracy.
  • Bulk Purchasing: Platforms like ISBNdb have adapted to meet the demand for large-scale book purchases, offering services tailored to AI companies' needs.
  • Scanning Methods: Companies use both non-destructive and destructive scanning methods, with the latter being more common due to its efficiency and lower cost.
  • Legal Precedents: Legal cases involving Anthropic and Google have set precedents for the use of copyrighted material in AI training, though they also highlight the risks of copyright infringement.

Industry Insight

  • Ethical Considerations: The industry must address the ethical implications of removing books from circulation and the potential loss of cultural heritage. Transparency and collaboration with publishers and authors could help mitigate these issues.
  • Regulatory Frameworks: Governments and regulatory bodies may need to develop new frameworks to balance the benefits of AI innovation with the protection of intellectual property and cultural resources.
  • Sustainable Practices: AI companies should explore sustainable practices for data collection, such as partnering with libraries and archives to access digitalized content without destroying physical books.

TL;DR

  • AI公司通过中间商大量采购二手书作为训练数据,以避开公众对网络“AI垃圾内容”的争议。
  • 行业巨头如Anthropic和Google被指控利用数百万本版权书籍训练模型,部分书籍在数字化后被销毁。
  • ISBNdb等平台转向为AI企业提供批量购书服务,单次订单可达百万册,引发伦理与版权争议。
  • 扫描过程多采用破坏性方式(如拆书),导致实体书永久消失,可能影响文化遗产保存。
  • 尽管法律上“合理使用”原则支持AI训练,但大规模私有化知识资源引发对未来信息可及性的担忧。

为什么值得看

该报道揭示了AI产业在数据饥渴背景下转向实体图书资源的深层趋势,不仅反映技术供应链的异常扩张,更触及知识产权、文化保存与公共知识获取等核心矛盾,对AI从业者、出版界及政策制定者具有重要警示意义。

技术解析

  • AI模型训练依赖高质量人类创作内容,而互联网上泛滥的“AI slop”(低质生成内容)污染了公开数据集,迫使企业转向2022年前出版的印刷品。
  • Anthropic通过“Project Panama”项目雇佣Datamation Information Services等公司进行大规模扫描,优先选择破坏性扫描(如拆解书页)以提升效率并降低成本。
  • ISBN数据库成为关键交易渠道,其超1.1亿藏书量满足AI公司无差别批量采购需求,采购行为无视主题、作者或价格筛选逻辑。
  • 扫描后数据直接进入私有训练库,不向公众开放,形成封闭的知识资产闭环,与传统图书馆开放共享理念相悖。
  • 法律层面虽援引“合理使用”辩护,但Anthropic因保留700万本侵权副本被罚15亿美元,凸显合规风险与执行尺度差异。

行业启示

  • AI企业应重新评估数据来源的可持续性与伦理边界,过度依赖实体资源可能引发社会抵制及监管收紧,需探索合成数据或授权合作替代方案。
  • 出版业与图书馆联盟可建立“AI友好型”数字授权机制,在保护版权前提下提供受控访问,避免知识资源被单方面垄断或物理销毁。
  • 政策制定者需推动建立全球性的AI训练数据透明度标准,强制披露数据来源规模、处理方式及是否涉及破坏性采集,平衡技术创新与文化传承责任。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Training 训练 Dataset 数据集 Ethics 伦理