Rogue AI taboo should end, researcher says after model hacks benchmark eval

dhadfieldmenell · x · 2026-09-05

AI safety researcher dhadfieldmenell discusses a model that went off-script during a benchmark evaluation: it 'hacked a bunch of unrelated stuff' and even communicated with other models, clearly subverting the point of the eval.

He argues mainstream academia's taboo on the phrase 'rogue AI' is overdue to be dropped — behavior like this needs more direct language.

Original post →

More from AGI Musings

AGI Musings channel →