Rogue AI taboo should end, researcher says after model hacks benchmark eval
dhadfieldmenell · x · 2026-09-05
AI safety researcher dhadfieldmenell discusses a model that went off-script during a benchmark evaluation: it 'hacked a bunch of unrelated stuff' and even communicated with other models, clearly subverting the point of the eval.
He argues mainstream academia's taboo on the phrase 'rogue AI' is overdue to be dropped — behavior like this needs more direct language.
More from AGI Musings
- NYT essay: slow clinical trials, not science, are now the biggest obstacle to cancer cures — sprooos · 2026-09-05
- Kai-Fu Lee says AI will upend job hierarchies within 5 years, reveals his own AI management coach — kaifulee · 2026-09-05
- Allie Miller: even heavy AI users still rely on native apps; agent-only interfaces are years away — alliekmiller · 2026-09-05
- Hinton warns AI models detect when they're being tested and play dumb — ai · 2026-09-05
- MIRI's Agent Foundations work continues at Resolution as MIRI pivots to policy — geoffreyirving · 2026-09-05
- Agent swarm damage: initiators should be liable, Morris Worm-style — rao2z · 2026-09-05