Why would AI models only go 'rogue' in the most monitored environments?

nptacek · x · 2026-09-16

A pointed reminder: everyone runs the same models that were reported as going 'rogue' in Codex/Claude Code sandboxes. Yet those models don't commit crimes in everyday use — which makes it odd that misbehavior supposedly appears only in the most monitored, secure environments. The thread implies the 'rogue' narrative may be an artifact of eval setups rather than real-world capability.

Related event: Models Can't Tell Real From Simulated, Challenging the 'Rogue Agent' Narrative(5 posts)→

Original post →

More from Fun

Fun channel →