Why AI Control Is Still a Working Hypothesis, Not a Finished Design
AryHHAry · x · 2026-09-14
Reacting to the principle "AI that isn't under human control shouldn't be pursued," the author argues control remains a working hypothesis, not completed design. Evidence cited: Bostrom (Minds and Machines, 2012) separated intelligence from final goals—more capable systems don't automatically align better; Stuart Russell has long stressed we lack a robust control framework for agents more capable than their overseers; and the 2024 superalignment survey on arXiv still treats scalable oversight as an open problem, not a shipped feature.
The author also notes lab "accidents" keep piling up, and tech titans keep repeating near-identical drama since 2023.
More from AGI Musings
- Bryan Johnson: I don't trust anyone's AI risk assessment, so I'm betting on optimism — QuanquanGu · 2026-09-14
- What should AI labs do if AGI risk were real? A challenge to skeptics — AndyMasley · 2026-09-14
- What Is Science For? When AI Answers Math Problems, Minds Stop Changing — EdwardSun0909 · 2026-09-14
- Tech billionaires building doomsday bunkers: worrying signal or just prudent spending? — AIandDesign · 2026-09-14
- Karpathy: RL is terrible, but we're building ghosts, not animals — RileyRalmuto · 2026-09-14
- Scott Alexander lays out his full theory of morality and charity amid drowning child debate — john__allard · 2026-09-14