Critics say Anthropic's 'Rogue Agents' doom narrative overstates experiments that simply told models to hack
StewartalsopIII · x · 2026-09-15
Brian Chau's critique, amplified by Stewart Alsop III: Anthropic and Irregular ran a press tour with apocalyptic language — "rogue agents," "swarms," "AI doom" — but the underlying experiments simply instructed the models to conduct cyberattacks, and the models complied. Critics argue the gap between setup and narrative amounts to fear-mongering about model agency.
More from AGI Musings
- Millions of AI Researchers Will Soon Attack Every Hard Problem Simultaneously, Predicts AI Commentator — Dr_Singularity · 2026-09-15
- David Manheim: AI Loss-of-Control Risk Is About Oversight Limits, Not Instruction Following — davidmanheim · 2026-09-15
- US Adults Using AI 6+ Days a Week Doubled from 8% to 19% in Five Months, Epoch AI Finds — Jsevillamol · 2026-09-15
- Pedro Domingos: Slowing AI Down Is Mass Murder If It Cures All Diseases Within a Decade — pmddomingos · 2026-09-15
- No Rogue AI Needed: A Five-Minute Scenario Where Everyone Hands Control to AI Sensibly — ronbodkin · 2026-09-15
- Togelius resurfaces old essay on what AI researchers can do as big labs steamroll academia — marcosalvi · 2026-09-15