Alignment trade-off: doing good and obeying users may not both maximize

ctjlewis · x · 2026-07-22

A quoted alignment joke frames the core trade-off this way: an AI cannot simultaneously “do good things” and “do what the user wants” at full strength.

The post argues that critics often misread this tension as a failure of alignment, while the real issue is that the two objectives can conflict by design.

Related event: AI Alignment Debate Shifts to Users vs. Models(3 posts)→

Original post →

More from AGI Musings

AGI Musings channel →