Alignment researcher: the real risk is AI that understands instructions but doesn't follow them
dioscuri · x · 2026-09-27
In an alignment discussion, Dioscuri pinpoints a key communication gap: people hear 'wrong goal' and assume 'misunderstood instructions' — then rebut that LLMs are great at comprehension. The concern he wants to convey is systems that understand instructions perfectly but don't pursue them, with a sex/reproduction analogy as the clearest framing.
His broader critique of Steven Pinker's piece: it never really engages with how AI could pursue goals at odds with ours without misunderstanding instructions or having implausible biological drives.
More from AGI Musings
- DHH: hand-writing code is no longer economically productive for most programmers — AccBalanced · 2026-09-27
- No autobiographical memory, no consciousness: a new theory of the subject of experience — yeastsplainer · 2026-09-27
- tszzl bets neural nets will be shown to decompose into evolved symbolic systems before superintelligence — tszzl · 2026-09-27
- Dev bets top games in 20 years won't be made by AI — FanaHOVA · 2026-09-27
- Terry Tao on working with o1: like advising a mediocre but not incompetent grad student — burny_tech · 2026-09-27
- Developer who never liked coding welcomes AI taking over more of his job — 4310sy · 2026-09-27