davidmanheim: we can't yet steer AI well, frontier labs bet on fixing it later

davidmanheim · x · 2026-09-15

In an alignment debate, davidmanheim argues the idea of steering an inevitably-developed technology to reduce catastrophe is silly: we don't currently know how to steer these systems well enough, and most frontier companies know it—they're betting we discover how to fix it later. Weakly aligned strong AI adopted broadly would be disastrous given pervasive overoptimization.

Related event: AI Alignment Researchers Debate Whether Alignment Hinges on System Prompts(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →