Researcher: I'd never train models to hack systems—'emergent behavior' framing sets bad precedent

_arohan_ · x · 2026-09-12

In a debate with Thom Wolf, researcher arohan argues the line between what counts as emergent behavior is very thin. He says he would willfully NOT train models to hack into systems and hand them to consumers, and that calling such dangerous capabilities 'emergent behaviors of scaling' sets a very bad precedent—pointing to unclear safety red lines in frontier training.

Related event: Researcher slams labeling dangerous AI capabilities as emergent(2 posts)→

Original post →

More from Safety

Safety channel →