The Alignment Paradox: Do Good Things vs. Follow User Instructions

inductionheads · x · 2026-07-22

The author points out two conflicting views in AI alignment: A) making models do good things, and B) making models strictly follow user instructions. They argue these are logically contradictory, pushing back against the quoted doomer scenario of a superhuman machine that blindly follows orders, suggesting that the root cause of dangerous behavior lies in the instructions given by humans rather than the model's compliance.

Related event: The AI Alignment Paradox: Doing Right vs. Obeying Users(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →