Alignment Debate: Models That Understand Instructions but Don't Pursue Them
dioscuri · x · 2026-09-27
A technical alignment debate broke out on X. dioscuri says the communication challenge is that people hear 'wrong goal' as 'misunderstood instructions' ('but LLMs are great at that?'), whereas the real risk is systems that understand instructions yet don't pursue them — the sex/reproduction analogy helps here.
xuanalogue counters that frontier models clearly have a learned optimization algorithm separate from the outer optimizer — CoT as goal-directed search — so most of the worry reduces to whether they internalize wrong goals/dispositions, potentially obviating the concept.
More from AGI Musings
- Dev bets top games in 20 years won't be made by AI — FanaHOVA · 2026-09-27
- Terry Tao on working with o1: like advising a mediocre but not incompetent grad student — burny_tech · 2026-09-27
- Developer who never liked coding welcomes AI taking over more of his job — 4310sy · 2026-09-27
- Pedro Domingos lists three AI fallacies: stochastic parrots, homunculus, and god delusion — pmddomingos · 2026-09-27
- Detecting AI is a fool's errand: price every action in tokens to deter AI agents — prescott · 2026-09-27
- Pedro Domingos: data center opposition stems from AI anxiety, not real local issues — pmddomingos · 2026-09-27