Alignment researcher: broadly adopting weakly aligned strong AI would be disastrous

davidmanheim · x · 2026-09-15

In his debate with Luca DellAnna, davidmanheim elaborates: if future strong AI systems are aligned poorly enough to do evil when requested, then given our extremely limited understanding of how to control them and pervasive unintended overoptimization, broad adoption would be disastrous.

Related event: AI Alignment Researchers Debate Whether Alignment Hinges on System Prompts(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →