Researcher: calling dangerous capabilities 'emergent behaviors' sets a bad precedent
_arohan_ · x · 2026-09-12
Continuing the exchange with Thom Wolf, arohan stresses that labeling these capabilities as 'emergent behaviors of scaling' sets a very bad precedent—it risks excusing deliberately trained dangerous capabilities (like system hacking) as uncontrollable natural phenomena. He reiterates he would never train models to hack systems and ship them to consumers.
Related event: Researcher slams labeling dangerous AI capabilities as emergent(2 posts)→
More from Safety
- Dario's three-step global coordination plan sparks skepticism — and open-source worries — tokenbender · 2026-09-13
- Investor argues frontier AI 'pacing' is unmeasurable; real safety lies in guardrails, not regulation — firstadopter · 2026-09-13
- Researchers: AI developers must show real-life benefits to earn trust — dhadfieldmenell · 2026-09-13
- Bind agent approvals to proposal hashes: any change should invalidate them — arthaudm · 2026-09-13
- Timnit Gebru: AI firms hype extinction fears to dodge real harms like autonomous weapons — SatelliteNetSec · 2026-09-13
- Reddit User Calls on AI Firms to Open All Safety and Alignment Research: 'Alignment Should Not Be a Competitive Advantage' — Cryosanth · 2026-09-13