Capability errors degrade smoothly, goal errors can be arbitrarily bad
gleech · x · 2026-09-07
Another segment of gleech's thread, arguing the live hypothesis that "capabilities are easier than alignment":
- Self-correction asymmetry: A coherent agent corrects its own mistaken beliefs because they hurt its goals, and protects its goals because revising them usually harms them.
- Fragility: Small practical errors degrade outcomes smoothly, but small errors in goals can be arbitrarily bad.
More from AGI Musings
- Computational Journalism: how interactive simulations could fix public debate — anselm · 2026-09-07
- dhh on AI Coding: Both Skeptics and Believers Are Right—Update Your Priors — bendee983 · 2026-09-07
- OpenAI Chief Scientist Jakub Pachocki: We Will See Machines Smarter Than Humans in Our Lifetime — oran_ge · 2026-09-07
- Why AI labs will close up: 'genie' pricing, hidden agent traces, and the four-minute-mile advantage — curious_vii · 2026-09-07
- AI is a competitive market, so surplus accrues to users, not vendors: Afinetheorem — Afinetheorem · 2026-09-07
- Redditor suspects flood of 'OpenAI achieved AGI' posts is a coordinated PR push — so_schmuck · 2026-09-07