Are frontier models' inner optimizers obsolete? CoT as goal-directed search sparks alignment debate
xuanalogue · x · 2026-09-27
xuanalogue and @dioscuri debate the "learned optimizer" concept in AI safety.
- xuanalogue argues frontier models already have a learned optimization algorithm beyond the outer optimizer — CoT as goal-directed search — making the inner-optimizer framing outdated; the real worry is internalized wrong goals/dispositions
- @dioscuri counters that the concept faded partly because it's irrelevant to recent attacks, "goal misgeneralization" is more precise for some reward hacking cases, and our understanding of model dispositions is now richer than "a learned optimizer inside"
More from AGI Musings
- 24 senior mathematicians at Harvard summit: PhDs should no longer hinge on dissertations — alejandroll10 · 2026-09-27
- GPT-6 Astra reportedly proves weak Goldbach case unconditionally in 2 pages — yshan2u · 2026-09-27
- Alignment researcher: the real risk is AI that understands instructions but doesn't follow them — dioscuri · 2026-09-27
- Why Would Genuinely Good Machine Gods Tolerate Despotic Regimes? — xuanalogue · 2026-09-27
- Dean Ball: AI safety and e/acc are natural allies — drop the kayfabe — deanwball · 2026-09-27
- Auto-research agents are coming, and peer review may shatter without defenses — askerlee · 2026-09-27