AirTag reveals how Amazon destroys rare books for AI training
Amazon's VGT3 team in Las Vegas systematically buys printed books, cuts off their spines for faster scanning, and destroys the originals to train its Nova AI models Anthropic conducted a similar operation called "Project Panama," buying books on marketplaces, spine-cutting them, and digitizing the content A judge ruled both companies' scanning practices qualified as fair use since originals were destroyed and not copied/resold AI companies are systematically scanning books by ISBN to access prin
Analysis
TL;DR
- Amazon's VGT3 team in Las Vegas systematically buys printed books, cuts off their spines for faster scanning, and destroys the originals to train its Nova AI models
- Anthropic conducted a similar operation called "Project Panama," buying books on marketplaces, spine-cutting them, and digitizing the content
- A judge ruled both companies' scanning practices qualified as fair use since originals were destroyed and not copied/resold
- AI companies are systematically scanning books by ISBN to access printed texts that predate 2022 and are free of AI-generated contamination
- The practice raises ethical concerns about destroying irreplaceable rare books and locking knowledge inside closed corporate AI models
Why It Matters
This reveals a hidden supply chain in the AI industry where physical cultural artifacts are being consumed as training data, raising urgent questions about intellectual property, fair use, and the environmental/cultural cost of AI development. For AI practitioners and researchers, it highlights the growing competition for high-quality, pre-2022 text data and the legal gray areas surrounding bulk digitization of copyrighted materials.
Technical Details
- Amazon's VGT3 team uses spine-cutting as a mechanical optimization to accelerate book scanning, indicating a high-throughput digitization pipeline designed for volume over preservation
- The scanning operation targets books by ISBN number, suggesting a systematic, database-driven approach to data collection rather than opportunistic acquisition
- Anthropic's "Project Panama" followed an identical spine-removal and digitization methodology, indicating an industry-standard practice for bulk book scanning
- Both companies use the scanned text to train proprietary language models (Amazon's Nova models, Anthropic's Claude models), converting physical books into private training corpora
- The legal framework currently permits this practice under fair use doctrine, as courts have ruled that destruction of originals rather than redistribution qualifies as transformative use
Industry Insight
- Expect increased regulatory scrutiny and potential legislation around bulk digitization of copyrighted materials, as the current fair use precedent may not survive challenges from publishing industries
- The race for clean, pre-2022 training data will likely intensify as more publishers and authors become aware of these practices, potentially driving up book acquisition costs or triggering boycotts
- AI companies should consider developing transparent data sourcing policies and exploring alternative training data strategies, as the current practice risks significant reputational damage and legal liability if it becomes widely known
Disclaimer: The above content is generated by AI and is for reference only.