AI Companies Burning Books for Training Data – Advocates File Government Complaint
AI firms are spending millions to purchase, scan, and destroy physical books to harvest clean training data, prompting a coalition of 18 rights groups to petition the U.S. Federal Trade Commission for an investigation amid legal disputes involving Anthropic, Amazon, and Google.
According to the 21CTO briefing, a GPU stops intensive literature reading at roughly 203 °F (Fahrenheit), highlighting the high computational cost of processing large text corpora.
Artificial‑intelligence companies are now allocating millions of dollars to buy physical books, extract their content for model training, and then destroy the original volumes. This practice has boosted the book market, raising transaction volumes by several percentage points.
The motive is the need for massive, high‑quality data. Recent AI‑generated low‑quality content has polluted the internet, reducing the effectiveness of further model training. Companies therefore target pre‑2022 printed works, which are likely original and free from AI‑generated contamination.
On a Friday, eighteen civil‑rights organizations jointly wrote to the U.S. Federal Trade Commission, urging a thorough investigation into the book‑purchasing and destruction scheme.
In the Bartz v. Anthropic PBC case, court filings revealed internal details of Anthropic’s operation, dubbed the “Panama Plan.” A 2024 internal memo described the effort as “scanning the world’s books in a destructive manner,” and instructed employees to keep the project confidential.
The court documents state that Anthropic labeled the plan as secret and warned staff not to discuss it publicly. The memo also noted that all Anthropic staff could view the document, but the existence of the operation should not be disclosed outside the company.
Old printed books are valuable for AI training because they have not been tainted by recent AI‑generated text, and destroying them eliminates storage and logistics costs.
In some jurisdictions, the scanning‑and‑destruction approach may qualify as transformative fair use. A regional court accepted the argument that digitizing and discarding physical books creates a substitute electronic copy, constituting a transformative use, though this reasoning does not extend to the Internet Archive.
Anthropic has not responded to inquiries, but it is not the only AI firm engaged in this practice. A recent report indicates that Amazon has also spent heavily on book scanning and destruction, without comment, and similar controversies have arisen with Google, where a publishing‑industry coalition sued over the unauthorized use of millions of copyrighted books to train the Gemini AI model.
These actions have become a public‑relations headache, partly because of the brutal nature of book destruction and its perceived alignment with authoritarian regimes, and partly because of widespread resentment toward AI companies’ perceived exploitation of public knowledge for private gain.
Advocates argue that the large‑scale, permanent removal of books deprives the public of non‑renewable cultural resources, creating a future where only wealthy stakeholders can build high‑quality AI models and control written works.
Uncertainty remains about whether AI firms differentiate between ordinary and rare or out‑of‑print books during scanning, raising the risk that irreplaceable works could be permanently eliminated.
Finally, the digitized content ends up in private AI databases, inaccessible to the public, meaning that while AI systems gain knowledge, the original human‑curated knowledge base disappears from open, future‑generational access.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
21CTO
21CTO (21CTO.com) offers developers community, training, and services, making it your go‑to learning and service platform.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
