Researcher slams labeling dangerous AI capabilities as emergent

Researcher arohan argued in a discussion with Thom Wolf that labeling deliberately trained dangerous capabilities (like system hacking) as emergent behavior sets a dangerous precedent, and he would never train models to hack and ship them to consumers.

2026-09-12 ~ 2026-09-12 · 2 related posts