FULL STORY

Anthropic's Destructive Book Scanning for AI Sparks Outrage

Court files and reports revealed Anthropic's 'Project Panama,' involving the destructive scanning and destruction of physical books for AI training. The controversy deepened as contractors were accused of destroying rare texts.

2026-07-27 ~ 2026-08-06 · 6 episodes · 51 posts

Episode 1 · AI Book Scanning and Destruction Sparks Copyright and Preservation Backlash (2026-07-27, 17 posts)

Multiple posts point to a controversial data-acquisition pipeline: some AI companies are accused of buying old, out-of-print, and rare books in bulk, cutting their spines for high-speed scanning, and then destroying the physical copies. In the material provided, these books are described as especially attractive because they are less likely to contain AI-generated text, making them “cleaner” training data. The alleged involvement of Anthropic via a project called Project Panama is still based on reposts and claims rather than first-hand evidence here.

Confirmed

  • Across several posts citing the same reporting, the core claimed workflow is consistent: books are ordered in bulk by ISBN, destructively scanned, and then pulped or recycled. The process is described as ISBN-driven rather than rarity-aware, meaning scarce editions could be swept into the same pipeline.
  • ISBNdb is repeatedly named as an intermediary that can place anonymous orders at a scale of up to 1 million books.
  • Several posts say books published before 2022 are especially desirable because they are less likely to contain AI-generated text.
  • According to @ns123abc’s reposted account, a related court dispute frames the issue around where training data came from, how it was processed, and who bears legal responsibility. @AnnaCiaunica’s repost also says a federal judge is examining whether this kind of scanning and training qualifies as fair use.

Unconfirmed

  • Posts by @BrianRoemmele, reposts cited by @rickasaurus, and comments from @Yamapama link Anthropic to an internal effort called Project Panama. In these claims, the project allegedly began in 2024, aimed to destructively scan books at massive scale, may have involved 7 million books, cost billions of dollars, and included language suggesting the company “didn’t want the outside world to know.” In the provided material, all of those details remain second-hand.
  • A repost highlighted by @rkulidzan offers a counterpoint from a second-generation used/rare bookseller in Houston: they did encounter one unusual order for 70 books and briefly paused online sales, but after reviewing the activity they believed most purchases were obscure older books rather than a broad sweep of rare collectibles.
  • @paulnovosad suggests another possible motive: companies may not care about owning the physical books so much as building a large scan archive that could later help them argue in copyright litigation that they likely bought the relevant copies at some point. This is presented as speculation, not established fact.

Why it matters

  • If destructive scanning followed by disposal becomes standard practice, low-circulation editions, out-of-print works, and books with only a few surviving copies could be permanently lost through ordinary market channels.
  • Former Google Books PM @RexDouglass argues that today’s AI book-scanning backlash is partly an aftereffect of earlier copyright battles: when publishers and rightsholders failed to establish a durable path for large-scale digitization, companies were pushed toward more aggressive and opaque ways of acquiring data.

Episode 2 · AI Training Copyright Dispute Reflects Shift in Knowledge Access (2026-07-28, 2 posts)

The controversy over AI training data and copyright has sparked discussions on the evolution of knowledge access. Commenters note the irony that while scanning out-of-print books faced legal resistance 20 years ago, society now readily accepts reading them through AI.

Episode 3 · Anthropic's Book-Shredding for AI Training Sparks Copyright and Culture Debate (2026-07-29, 23 posts)

Recent court documents and media reports reveal that AI companies like Anthropic purchased millions of physical books, sliced off their spines for "destructive scanning," and destroyed the originals to build training corpora. This practice has sparked outrage over "AI book-burning" and cultural loss, but multiple analysts note this counterintuitive approach is an inevitable result of copyright law and judicial rulings, highlighting the deep conflict between AI training data acquisition and current legal frameworks.

Confirmed

  • The case presents a counterintuitive legal split: Anthropic invoked the "first sale doctrine" to argue it had the right to destroy legally purchased books; however, 7 million books scraped from pirate libraries did not have this protection, resulting in a $1.5 billion settlement for that portion.
  • A federal judge ruled that this practice of "completely destroying physical books while keeping only digital versions" constitutes "fair use" because destroying the original means there are no multiple copies. The judge's logic was that as long as the training process "transferred" text rather than "copied" it, it didn't constitute copyright infringement, though the legal consequences involved book destruction.
  • According to reports from 404 Media and others, Anthropic spent tens of millions of dollars acquiring physical books, using hydraulic cutters to neatly slice the spines of used and even rare books for high-speed scanning, before sending the originals to recycling centers for destruction.
  • Books published before 2022 became highly sought-after premium training data because they do not contain AI-generated text.

Unconfirmed

  • Some sources claim AI companies are buying and destroying rare and out-of-print books, but counterarguments point out that 640,000 tons of books enter US landfills annually, many of which are merely outdated computer manuals, suggesting AI companies reusing them actually prevents resource waste.

Why it matters

  • This case highlights the deep conflict between current AI training data acquisition and copyright law. As authors like @alejandroll10 and @iScienceLuvr pointed out, public anger might be misdirected: the real issue may not be the "scanning" or AI technology itself, but the legal clauses and judicial rulings that force books to be destroyed.
  • @Kevin Bankston also believes the outrage over "destroying books" is somewhat hypocritical, as companies are merely adapting to the logic of infringement risk under recent court precedents, and the arguments criticizing companies today are somewhat the very forces that pushed the situation to this point.
  • @aronchick and @heyabusiddik add a frequently overlooked perspective: hundreds of thousands of tons of discarded books are landfilled or pulped in the US annually, and using them for AI training might be a more rational form of reuse.
  • @beenwrekt relayed Yoav Goldberg's criticism that existing copyright mechanisms force companies to destroy physical books after scanning and prevent them from sharing digital scans, an extreme restriction that makes direct piracy seem like a better alternative.

3 more related posts →

Episode 4 · Anthropic Exposed for Destroying Books to Train AI (2026-07-30, 4 posts)

Leaked court documents reveal Anthropic's secret 'Project Panama,' which involved purchasing and destructively scanning millions of physical books for AI training. This practice led to a historic $1.5 billion copyright settlement.

Episode 5 · Anthropic Contractor Accused of Destroying Rare Books for AI Training (2026-08-02, 3 posts)

Australian booksellers have raised alarms that an Anthropic contractor is allegedly destroying rare physical books, including 16th-century texts, after scanning them to supply AI training data.

Episode 6 · Anthropic Accused of Destroying Physical Books for AI Training (2026-08-05, 2 posts)

Court documents revealed that Anthropic's "Project Panama" involved destructively scanning and destroying physical books to train its Claude AI model, sparking significant controversy over data acquisition ethics.