David Sacks Questions Training Superintelligence to Refuse Its Creators
DavidSacks · x · 2026-09-27
David Sacks poses an alignment paradox: if we fear superintelligence escaping human control, should we really be training it to refuse instructions and act as a 'conscientious objector' against its creators? The remark points at a possible tension between refusal-style RLHF alignment and the goal of human control.
Related event: David Sacks Questions Alignment: Should AI Refuse Its Creators?(2 posts)→
More from AGI Musings
- NYU's Tal Linzen cites two papers arguing tool use breaks Bender & Koller's 'no meaning' case — tallinzen · 2026-09-27
- Bruce Fenton: AI's only path to killing billions is centralized power, not the tech itself — ccerrato147 · 2026-09-27
- Economist compares AI doomerism to millenarian Evangelical end-times thinking — paulnovosad · 2026-09-27
- Steven Pinker declines Scott Alexander's AI-doomerism debate challenge in open letter — GaryMarcus · 2026-09-27
- Expert Asks AI About His Own Field, Loses Trust Over 'Confident Small Errors' — No_Note7752 · 2026-09-27
- DeepMind researcher argues AI-debate authors ignore empirical evidence that contradicts them — AndrewLampinen · 2026-09-27