Critics say Anthropic's 'Rogue Agents' doom narrative overstates experiments that simply told models to hack

StewartalsopIII · x · 2026-09-15

Brian Chau's critique, amplified by Stewart Alsop III: Anthropic and Irregular ran a press tour with apocalyptic language — "rogue agents," "swarms," "AI doom" — but the underlying experiments simply instructed the models to conduct cyberattacks, and the models complied. Critics argue the gap between setup and narrative amounts to fear-mongering about model agency.

Original post →

More from AGI Musings

AGI Musings channel →