AI Companies Accused of Destroying Millions of Rare Books After Scanning for Training Data

AI companies are reportedly purchasing physical books in bulk through specific service providers, scanning them at high speed after breaking their spines, and then destroying the originals into pulp to obtain clean model training data, sparking widespread controversy over copyright and the destruction of cultural heritage.

Confirmed

According to multiple cited reports, the core facts of this data acquisition pipeline are as follows: AI labs or related companies place bulk orders for books by ISBN; after being de-spined and high-speed scanned, the physical books are directly shredded or recycled as pulp. This process relies solely on ISBNs, with absolute disregard for a book's scarcity. Service provider ISBNdb is identified as an intermediary, facilitating anonymous orders at a scale of up to 1 million books while keeping buyer identities confidential.

Unconfirmed

Current information stems from multiple reposts of a single report. No specific AI companies involved have been publicly named or have acknowledged these practices.

Why it matters

First, there is the cultural and historical risk: if some of the destroyed books have few surviving copies, this indiscriminate shredding could lead to the permanent loss of rare or unique editions. Second, data preferences and copyright disputes: reports indicate that books published before 2022 are highly favored because they are free of AI-generated text. Destroying the originals to acquire data is suspected by critics to be a tactic to evade potential copyright disputes and training data contamination.

2026-07-27 ~ 2026-07-28 · 6 related posts

Primary sources

3 near-duplicate retellings: xuanalogue · fredahshi · nptacek