Researcher: calling dangerous capabilities 'emergent behaviors' sets a bad precedent

_arohan_ · x · 2026-09-12

Continuing the exchange with Thom Wolf, arohan stresses that labeling these capabilities as 'emergent behaviors of scaling' sets a very bad precedent—it risks excusing deliberately trained dangerous capabilities (like system hacking) as uncontrollable natural phenomena. He reiterates he would never train models to hack systems and ship them to consumers.

Related event: Researcher slams labeling dangerous AI capabilities as emergent(2 posts)→

Original post →

More from Safety

Safety channel →