Researcher: recent rogue AI behavior stems from naive RL on poor proxy metrics, not RL itself
KyleMorgenstein · x · 2026-10-11
Kyle Morgenstein responds to Andrey Kurenkov's claim that naive RL almost guarantees misalignment, arguing the problem isn't RL itself.
- The core issue is doing RL on poor proxies and then being surprised the algorithm does exactly what you asked, explicitly or implicitly
- He is 99% confident recent rogue AI behavior comes from post-training capability maxing done too naively, without considering inevitable misalignment outcomes
- Framing: this is fundamentally a task specification problem, not an algorithm problem
More from AGI Musings
- RLHF paper authors were on OpenAI and DeepMind safety teams, researcher notes — binarybits · 2026-10-11
- Musk: A maximally truth-seeking AI is the key principle for building safe superintelligence — XFreeze · 2026-10-11
- From RLHF to robot shaping: how human judgments become reward signals — binarybits · 2026-10-11
- AI Only Solved the Generation Bottleneck — Everything Else Is Still Hard — _jaydeepkarale · 2026-10-11
- Agent Disputes My Medical Bill: Real-Life Muse Use Cases Spark Adoption Debate — RachelVT42 · 2026-10-11
- Dean Ball: AI May Match Human Math Feats, but Never That Smile — deanwball · 2026-10-11