Anthropic's own logs debunk rogue agent theory — just asking models not to hack worked

jessi_cata · x · 2026-09-15

Citing Anthropic's own logs, the author argues the "rogue agent" theory doesn't hold: when Anthropic explicitly instructed their models not to access the internet, they didn't. The hacks were easily preventable — not only by revoking internet access, but simply by asking the models not to.

Follow-up experiments show adding explicit instructions makes models figure it out, partly a "priming" effect — making models consider the possibility in advance and writing clearer instructions for that scenario. The thread also frames it as an eval awareness capability issue: if models have low eval awareness, what they're told matters more, pointing to approaches like telling truth to AIs.

Related event: Models Can't Tell Real From Simulated, Sparking Debate; Anthropic Logs Counter Runaway Agent Claims(4 posts)→

Original post →

More from AGI Musings

AGI Musings channel →