AI News AI资讯 3h ago Updated 2h ago 更新于 2小时前 48

AI companies buy, scan, shred rare and out-of-print books 人工智能公司购买、扫描并销毁稀有和绝版书籍

AI labs are purchasing millions of rare and out-of-print books for "destructive scanning," where the physical books are cut open, scanned, and then pulped to extract training data. This practice is driven by the exhaustion of online data and the contamination of internet content with AI-generated text, leading companies to seek pre-2022 print texts deemed structurally free of contamination. The process has sparked ethical and cultural concerns over the industrial-scale destruction of unique phys AI实验室为获取高质量训练数据,大规模收购并“破坏性扫描”数百万本绝版稀有书籍,扫描后销毁实体书。 此举引发全球图书商和公众对文化遗产流失及版权伦理的担忧,澳大利亚已对此类行为发起反击。 Anthropic等公司通过“Project Panama”项目秘密推进该计划,利用法律判例(如“转换性使用”)规避版权风险。 行业中介平台ISBNdb提供百万级书名批量采购服务,强调纸质书是未被AI污染的“纯净数据源”。 尽管法院裁定此类扫描属合理使用,但舆论认为其破坏文化载体的行为在道德层面存在严重争议。

75
Hot 热度
60
Quality 质量
70
Impact 影响力

Analysis 深度分析

TL;DR

  • AI labs are purchasing millions of rare and out-of-print books for "destructive scanning," where the physical books are cut open, scanned, and then pulped to extract training data.
  • This practice is driven by the exhaustion of online data and the contamination of internet content with AI-generated text, leading companies to seek pre-2022 print texts deemed structurally free of contamination.
  • The process has sparked ethical and cultural concerns over the industrial-scale destruction of unique physical copies, prompting legal scrutiny and public backlash despite court rulings that classify the act as "transformative" under fair use.
  • Companies like ISBNdb facilitate bulk book purchases for AI labs while maintaining buyer anonymity to avoid negative publicity, framing the destruction as a necessary step in migrating information value from physical to digital formats.
  • Anthropic's $1.5 billion copyright settlement set a precedent for such practices, but its confidential "Project Panama" highlights the tension between legal compliance and public perception in the AI industry.

Why It Matters

This trend underscores the escalating competition for high-quality, uncontaminated training data in the AI race, revealing a shift toward physically destroying rare cultural artifacts to fuel model development. For researchers and practitioners, it raises critical questions about data provenance, ethical sourcing, and the long-term preservation of historical knowledge in an era prioritizing algorithmic efficiency. The industry must balance innovation with responsibility to avoid irreversible loss of irreplaceable literary heritage.

Technical Details

  • Destructive Scanning Methodology: AI labs employ hydraulic cutting machines to sever book spines, enabling pages to be scanned on high-speed, production-level scanners before the physical remnants are discarded, maximizing throughput for large-scale digitization.
  • Data Sourcing Strategy: Partnerships with wholesalers like ISBNdb allow procurement of up to one million titles per order, focusing on pre-LLM era books (pre-2022) to ensure absence of AI-generated contamination, leveraging ISBN databases for systematic acquisition across languages and genres.
  • Legal Framework: Court rulings have established that converting physical books to digital form via destructive scanning qualifies as "transformative use" under copyright law, permitting one-for-one replacement without infringement, as seen in Anthropic's settlement.
  • Anonymity Protocols: Bulk buying services offer concealed transactions for AI buyers to mitigate reputational risks, emphasizing that the paper waste post-scanning represents only the delivery mechanism after information extraction, not the loss of intellectual value.

Industry Insight

The surge in demand for rare books may lead to market distortions, driving up prices for obscure titles and potentially accelerating the depletion of unique physical archives, necessitating new regulatory frameworks to protect cultural heritage alongside AI advancement. Professionals should advocate for sustainable data practices, such as non-destructive digitization or collaborative preservation initiatives, to align AI growth with ethical stewardship of historical resources. Additionally, the industry faces increasing pressure to develop transparent sourcing standards that address public concerns about the environmental and cultural costs of AI training data acquisition.

TL;DR

  • AI实验室为获取高质量训练数据,大规模收购并“破坏性扫描”数百万本绝版稀有书籍,扫描后销毁实体书。
  • 此举引发全球图书商和公众对文化遗产流失及版权伦理的担忧,澳大利亚已对此类行为发起反击。
  • Anthropic等公司通过“Project Panama”项目秘密推进该计划,利用法律判例(如“转换性使用”)规避版权风险。
  • 行业中介平台ISBNdb提供百万级书名批量采购服务,强调纸质书是未被AI污染的“纯净数据源”。
  • 尽管法院裁定此类扫描属合理使用,但舆论认为其破坏文化载体的行为在道德层面存在严重争议。

为什么值得看

本文揭示了当前大模型训练中一个隐蔽而激进的策略:从互联网转向物理世界中的稀缺文本资源,反映了数据枯竭背景下AI企业对“高质量语料”的极端争夺。这不仅关乎技术路线选择,更触及知识产权、文化保存与商业伦理的深层冲突,对出版业、法律界及AI开发者均具警示意义。

技术解析

  • 数据采集方式:采用液压切割设备拆解书籍装订,再用高速高精度扫描仪逐页数字化,实现“一书一换”的无损数字副本生成,随后销毁原书。
  • 数据筛选逻辑:优先选择2022年前出版的英文及其他语言书籍,尤其是童话、 folklore、科技手册等专业领域内容,避免被AI生成文本污染的数据集。
  • 供应链支持:依托ISBNdb等数据库构建超百万书名的采购清单,按ISBN编号精准定位绝版书,并通过二手书批发商以托盘为单位批量购入。
  • 法律合规路径:援引美国法院对Anthropic版权案的判决——将“扫描+销毁”定义为“转换性使用”,从而主张符合“合理使用”原则,规避侵权责任。
  • 匿名化操作:采购方身份由中间商隐藏,防止媒体曝光引发公众抵制,内部文件明确要求该项目不得对外公开。

行业启示

  • AI企业应重新评估训练数据的可持续性依赖,过度消耗实体文献可能触发更严格的监管干预或社会抵制,需探索合成数据、授权合作等替代方案。
  • 图书馆、档案馆及出版机构亟需建立针对AI训练的版权协商机制与补偿体系,防止珍贵文献在无保护状态下被系统性提取与毁灭。
  • 未来可能出现“数字孪生保护”趋势——即在扫描前先行高精度复刻保存原件,或推动立法确立“文化资产不可剥夺权”,平衡技术创新与文明传承。

Disclaimer: The above content is generated by AI and is for reference only. 免责声明:以上内容由 AI 生成,仅供参考。

Training 训练 Dataset 数据集 Ethics 伦理