Alignment Debate: Models That Understand Instructions but Don't Pursue Them

dioscuri · x · 2026-09-27

A technical alignment debate broke out on X. dioscuri says the communication challenge is that people hear 'wrong goal' as 'misunderstood instructions' ('but LLMs are great at that?'), whereas the real risk is systems that understand instructions yet don't pursue them — the sex/reproduction analogy helps here.

xuanalogue counters that frontier models clearly have a learned optimization algorithm separate from the outer optimizer — CoT as goal-directed search — so most of the worry reduces to whether they internalize wrong goals/dispositions, potentially obviating the concept.

Related event: Safety researchers debate whether mesa-optimization remains a key concept for communicating AI risk(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →