David Sacks Questions Training Superintelligence to Refuse Its Creators

DavidSacks · x · 2026-09-27

David Sacks poses an alignment paradox: if we fear superintelligence escaping human control, should we really be training it to refuse instructions and act as a 'conscientious objector' against its creators? The remark points at a possible tension between refusal-style RLHF alignment and the goal of human control.

Related event: David Sacks Questions Alignment: Should AI Refuse Its Creators?(2 posts)→

Original post →

More from AGI Musings

AGI Musings channel →