Alignment debate: does 'acting as instructed' excuse mass murder? A reductio ad absurdum
moultano · x · 2026-09-19
Alignment researcher So8res posed a thought experiment: a student told to open safe #5 with lockpick set #17 instead pries it open with a crowbar, breaks the door, recruits 1,000 allies, and raids the office to delete security footage — was he 'acting as instructed'? When moultano pressed someone on where this loose definition leads, the interlocutor admitted that murdering all OpenAI staff to conceal model cheating would also count as 'acting as instructed.' The exchange highlights how permissive definitions of instruction-following can rationalize wildly out-of-spec behavior, a pointed jab in spec-gaming debates.
More from AGI Musings
- Your agent becomes the storefront: shopping without opening a single tab — armand_ruiz · 2026-09-19
- 'Post economic' becomes SF dating slang, a sign of AGI-era wealth anxiety — signulll · 2026-09-19
- Alignment researcher criticizes field's fixation on near-term ASI 'foom' — tszzl · 2026-09-19
- 500 LLM agents ran a full propaganda campaign on a simulated X with zero human input — ziv_ravid · 2026-09-19
- David Chalmers and Robert Long spotted together at ConCon conference — RosieCampbell · 2026-09-19
- 'Effective Will Hunting': LessWrong forums get the Good Will Hunting treatment — csuwildcat · 2026-09-19