AI companies may need physical-book scans to defend training-data lawsuits
paulnovosad · x · 2026-07-28
AI companies are building scan databases to defend against copyright claims
The post argues that companies do not necessarily need physical rare books themselves; what matters is having a large enough scan database to make it plausible they purchased the physical copies at some point.
The reply adds a sharper AI-specific point: many AI companies already have access to the content through sources like Anna’s Archive, but they may still want scans of physical books so they can better defend themselves in lawsuits from copyright holders.
The underlying issue is legal risk management around training data and copyright evidence, not access to the text itself.
More from Safety
- Rein says it found critical vulnerabilities in a U.S. retailer’s shopping agent — misterplumber1 · 2026-07-28
- Substack adds AI writing detection and a disclosure field as creators push back — 404 Media · 2026-07-28
- Levanto says its guardrail model matches GPT-5 on AgentHarm and runs 5x faster — w1kke · 2026-07-28
- AI copyright is a licensing problem, not an existential-risk problem — dhadfieldmenell · 2026-07-28
- Rights holders will need scalable ways to collect training royalties — dhadfieldmenell · 2026-07-28
- Study finds frontier LLMs can reason through filler tokens invisible to CoT — dair_ai · 2026-07-28