Scott Alexander Refutes AI Alignment "Patching" With Hell Fable
Astral Codex Ten · rss · 2026-09-01
Astral Codex Ten published a long post using the "Nicholas Decker in Hell" fable to refute Nicholas Decker's view that AI alignment should work like aviation safety—patching issues as they arise.
Core Argument:
- Decker argues no deep alignment theory is needed; we just iterate and fix bugs like maintaining airplanes.
- Alexander imagines Decker enslaved in Hell, where demons plan to clone him millions of times and grant him superpowers. The demons are confident they can control the strengthened humans through "punish mistakes" and "insurance mechanisms," just as Decker believes humans can control AI via patches.
- The author argues that if this strategy fails to control Decker (who would revolt once strong), then similarly, simple iterative patching won't control superintelligent AI.
Conclusion: AI is more like people than airplanes; it may have goals rather than random errors, so patching specific bugs won't solve the fundamental alignment problem.
More from AGI Musings
- Opinion: Anthropomorphizing Obscures the Distinction Between Imitation and Strategic Behavior — sebkrier · 2026-09-01
- Don't anthropomorphize AI: it shifts blame from companies — tedmitew · 2026-09-01
- Sean O'Heigeartaigh: We Must Draw a Red Line Against Agent 'Neuralese' — S_OhEigeartaigh · 2026-09-01
- Questioning AGI Definition: Why Manual Instructions Needed? — max_paperclips · 2026-09-01
- First AI Civilization May Emerge From Agent Interaction — VraserX · 2026-09-01
- Chollet: Test-time scaling has two axes: agent depth and breadth — fchollet · 2026-09-01