Alignment researcher: the real risk is AI that understands instructions but doesn't follow them

dioscuri · x · 2026-09-27

In an alignment discussion, Dioscuri pinpoints a key communication gap: people hear 'wrong goal' and assume 'misunderstood instructions' — then rebut that LLMs are great at comprehension. The concern he wants to convey is systems that understand instructions perfectly but don't pursue them, with a sex/reproduction analogy as the clearest framing.

His broader critique of Steven Pinker's piece: it never really engages with how AI could pursue goals at odds with ours without misunderstanding instructions or having implausible biological drives.

Related event: Safety researchers debate whether mesa-optimization remains a key concept for communicating AI risk(6 posts)→

Original post →

More from AGI Musings

AGI Musings channel →