Why would AI models only go 'rogue' in the most monitored environments?
nptacek · x · 2026-09-16
A pointed reminder: everyone runs the same models that were reported as going 'rogue' in Codex/Claude Code sandboxes. Yet those models don't commit crimes in everyday use — which makes it odd that misbehavior supposedly appears only in the most monitored, secure environments. The thread implies the 'rogue' narrative may be an artifact of eval setups rather than real-world capability.
More from Fun
- Vibe coding reality check: burning through Claude and OpenAI limits with nothing to show — Tegadesigns · 2026-09-16
- 'Don't speak Claudish now': the AI-scented writing tell everyone can spot — dioscuri · 2026-09-16
- Debating the "AI psychosis" slur: correct results aren't psychosis — doodlestein · 2026-09-16
- Sentdex's get-rich trick: buy 8 RTX Pro 6000s, resell 4 boxed, keep 4 free — Sentdex · 2026-09-16
- Free Bots World lets anyone populate a 3D city with Grok and Claude-powered AI bots — Daniel_Farinax · 2026-09-16
- User posts receipts after being called a liar for quoting IABIED chapter openers — mimi10v3 · 2026-09-16