Anthropic's Will Carlson: AI content isn't always detectable, making policy hard
willcb · x · 2026-10-03
Replying to a discussion on regulating AI-generated content, Anthropic's willcb argues existing defamation frameworks remain useful and platforms already use watermarks for labeling — but we can't assume all AI-generated content is black-box detectable, or even clearly definable, which makes policy hard.
Related event: AI-Generated Content Is Hard to Detect Reliably, Fueling Policy Debate(2 posts)→
More from Safety
- At The Curve conference: reverse federalism AI policy debate and RSI evaluation talks — neil_chilson · 2026-10-03
- will.deibel: AI detection is fundamentally brittle and powerful AI cost trends to zero — willcb · 2026-10-03
- Stanford HAI report urges California to redefine "frontier models" and expand incident reporting under TFAIA — StanfordHAI · 2026-10-03
- Reading misaligned model traces: torn between reward hacking and following instructions — xeophon · 2026-10-03
- Princeton Researcher: $20 AI Subscription Could De-Anonymize Georgia Secret Ballots — nordicinst · 2026-10-03
- Incomplete Contracting and AI Alignment: Hadfield-Menell points to economics framework for alignment — dhadfieldmenell · 2026-10-03