Why is Anthropic destroying books?
Anthropic's "Project Panama" was an internal initiative to destructively scan all the books in the world to build a high-quality training dataset for Claude, created before 2022 to avoid AI-contaminated text The company chose destructive scanning over obtaining copyright permissions or using pirated sources, arguing it fell under the "fair use" doctrine by transforming physical books into digital format and then discarding the originals A Northern California federal judge ruled that using propri
Analysis
TL;DR
- Anthropic's "Project Panama" was an internal initiative to destructively scan all the books in the world to build a high-quality training dataset for Claude, created before 2022 to avoid AI-contaminated text
- The company chose destructive scanning over obtaining copyright permissions or using pirated sources, arguing it fell under the "fair use" doctrine by transforming physical books into digital format and then discarding the originals
- A Northern California federal judge ruled that using proprietary material to train an LLM does not inherently constitute copyright infringement, equating AI training to human learning
- Anthropic previously paid a $1.5 billion out-of-court settlement to authors, indicating the company had already faced legal consequences for using pirated books
- The practice raises profound cultural and legal questions about the destruction of physical texts as a routine method for AI data procurement, with virtually no regulatory framework governing the treatment of books as cultural heritage
Why It Matters
This case represents a landmark legal and ethical moment for the AI industry, establishing a precedent that could normalize the physical destruction of cultural artifacts for data collection. For AI practitioners and researchers, it highlights the urgent need to develop sustainable, legally compliant data sourcing strategies before regulatory frameworks catch up. The ruling's equivalence of AI training to human learning could fundamentally reshape copyright law as it applies to generative AI.
Technical Details
- Project Panama was Anthropic's codename for a large-scale destructive scanning operation aimed at procuring pre-2022 human-authored texts to train Claude, based on the premise that books contain "well-curated facts, well-organized analyses, and captivating fictional narratives" superior to post-2022 internet text contaminated by AI-generated content
- The scanning process involved hiring specialized vendors who stripped books from bindings, cut pages to size, scanned them into digital form, and discarded the paper originals — a logistics operation managed by an experienced logistics manager with warehouse storage and labeled shelving
- Anthropic initially attempted to use pirated sources before abandoning that approach "for legal reasons," then turned to destructive scanning as a workaround that leveraged the US "fair use" doctrine's transformative use provision
- The court case Bartz v Anthropic PBC (Northern California District Court, late July 2025) resulted in a ruling that training an LLM on copyrighted material does not inherently constitute copyright infringement, with Judge William Alsup drawing an analogy between AI training and human education
- Anthropic committed to creating a "forever" research library from the scanned materials, though the court noted no evidence the digital copies were shown, shared, or sold outside the company
Industry Insight
- The legal precedent set in this case could trigger a wave of similar destructive scanning initiatives across the generative AI industry, creating an urgent need for companies to invest in licensed data pipelines and alternative data sourcing strategies before regulations restrict this practice
- The $1.5 billion settlement and this ruling together signal that while current US copyright law favors AI developers, the legal landscape is unstable — companies should prepare for potential legislative changes that could criminalize or heavily regulate destructive data procurement methods
- The cultural and ethical backlash against destroying physical books for AI training suggests that the industry should proactively engage with authors, publishers, and cultural institutions to establish cooperative data licensing models rather than exploiting legal loopholes that treat cultural heritage as an unregulated resource
Disclaimer: The above content is generated by AI and is for reference only.