FULL STORY

Anthropic Faces Backlash Over Destructive Book Scanning

Leaked court documents reveal Anthropic destructively scanned and destroyed millions of physical books to train Claude, sparking severe copyright and ethical backlash.

2026-07-28 ~ 2026-07-31 · 3 episodes · 27 posts

Episode 1 · AI Training Copyright Dispute Reflects Shift in Knowledge Access (2026-07-28, 2 posts)

The controversy over AI training data and copyright has sparked discussions on the evolution of knowledge access. Commenters note the irony that while scanning out-of-print books faced legal resistance 20 years ago, society now readily accepts reading them through AI.

Episode 2 · AI Firms Face Backlash Over Destroying Books for Training Data (2026-07-29, 23 posts)

Anthropic is currently embroiled in a fierce debate over the copyright of AI training data. The controversy centers on how the company handles physical books used for training: destroying them after scanning. While this practice has drawn intense public backlash, multiple analysts point out that it is actually an inevitable outcome shaped by copyright law and judicial rulings, rather than simply a malicious corporate choice.

Confirmed

  • The case reveals a counterintuitive legal distinction: Anthropic invoked the "first sale doctrine" to argue it had the right to destroy legally purchased books; however, the 7 million books scraped from pirated databases do not enjoy this protection, leading to a 15 亿美元 settlement for that portion of the lawsuit.
  • The judge's reasoning suggested that as long as the training process "transferred" text rather than "copied" it, it did not constitute copyright infringement, though the legal consequences ultimately involved the destruction of the books.

Unconfirmed

  • The specific application details and practical impact of the "superintelligent lawyer" thought experiment within this case.

Why It Matters

  • This case highlights the deep-rooted contradictions between current AI training data acquisition and copyright law. As authors like @alejandroll10 and @iScienceLuvr pointed out, public outrage might be somewhat misguided: the real issue may not be the "scanning" or the AI technology itself, but rather the legal clauses and judicial rulings that force books to be destroyed. @Kevin Bankston also believes that the外界 anger over "book destruction" is a bit melodramatic, as the company is merely adapting to the infringement risk logic established by court precedents. Ironically, the very arguments criticizing the company today are, to some extent, the same forces that helped drive this situation in the first place.

3 more related posts →

Episode 3 · Leaked Documents Reveal Anthropic's Destructive Book Scanning for AI Training (2026-07-30, 2 posts)

Leaked court documents reveal Anthropic secretly scanned millions of physical books to train Claude, leading to a massive $1.5 billion copyright settlement.