Commenter argues Anthropic's book training garbles texts, 'losing' millions of books

StewartalsopIII · x · 2026-09-14

Stewart Alsop III argues that Anthropic scanned millions of books to train its LLMs, but trained the model to avoid reproducing direct quotes for copyright reasons — with the side effect that the model garbles the original authors' wording. In his view, this effectively makes those books 'lost to civilization,' calling it 'the largest loss of knowledge since the dark ages.' The claim is opinionated commentary on the tension between copyright compliance and textual fidelity in LLM training; Anthropic has not confirmed the training details.

Related event: Investor Claims Anthropic Destroyed Millions of Books to Train LLMs(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →