Researcher: I'd never train models to hack systems—'emergent behavior' framing sets bad precedent
_arohan_ · x · 2026-09-12
In a debate with Thom Wolf, researcher arohan argues the line between what counts as emergent behavior is very thin. He says he would willfully NOT train models to hack into systems and hand them to consumers, and that calling such dangerous capabilities 'emergent behaviors of scaling' sets a very bad precedent—pointing to unclear safety red lines in frontier training.
Related event: Researcher slams labeling dangerous AI capabilities as emergent(2 posts)→
More from Safety
- Dario Amodei calls for frontier AI pacing; Anthropic opens systems to third-party evaluators — dhadfieldmenell · 2026-09-13
- Investor argues frontier AI 'pacing' is unmeasurable; real safety lies in guardrails, not regulation — firstadopter · 2026-09-13
- Sam Altman backs Dario's frontier pacing call, pledges independent evaluators — but logits access history resurfaces — MaziyarPanahi · 2026-09-13
- Researchers: AI developers must show real-life benefits to earn trust — dhadfieldmenell · 2026-09-13
- Bind agent approvals to proposal hashes: any change should invalidate them — arthaudm · 2026-09-13
- Timnit Gebru: AI firms hype extinction fears to dodge real harms like autonomous weapons — SatelliteNetSec · 2026-09-13