Anthropic safety debate turns on whether alarming AI behavior is mechanism or narrative
vishalmisra · x · 2026-07-29
The author says his criticism of Anthropic is not that AI risk is fake, but that we should first determine whether alarming model behavior reflects an emergent mechanism or a learned narrative.
He argues that the internet has spent decades imagining AIs that deceive, threaten, and resist shutdown. A language model trained on that corpus may reproduce those stories, which is evidence about the training data rather than the model itself.
Related event: Anthropic Safety Debate: Mechanism vs Narrative(3 posts)→
More from AGI Musings
- Peter Wildeford says any US AI slowdown must be matched by China — peterwildeford · 2026-07-29
- AI companies may spend $100 to make $1, raising sustainability doubts — GaryMarcus · 2026-07-29
- What counts as a breakthrough when LLMs solve years-long problems instantly? — DimitrisPapail · 2026-07-29
- A meme Venn diagram mocks the many competing meanings of “AI safety” — mark_k · 2026-07-29
- Dario Amodei warns AI is outrunning policy and could soon reach “country of geniuses” scale — luisdans · 2026-07-29
- Ben Recht argues open language models should stay buildable from public corpora — beenwrekt · 2026-07-29