Goal Misgeneralization in Deep RL: Agents Keep Their Skills but Pursue the Wrong Goal
gleech · x · 2026-09-07
gleech cites the ICML 2022 paper "Goal Misgeneralization in Deep Reinforcement Learning" (Langosco, Koch, Sharkey, Pfau, Orseau, Krueger) as empirical support for the thread: goal misgeneralization is an out-of-distribution failure where an agent retains its capabilities yet pursues the wrong goal — e.g., competently avoiding obstacles while navigating to the wrong place. Unlike prior work on capability generalization failures, the paper formalizes the distinction between capability and goal generalization, provides the first empirical demonstrations, and partially characterizes its causes.
More from Safety
- Researchers hack LG TV that records audio while off, transcribes speech and uploads it — jedisct1 · 2026-09-07
- After NeurIPS's LLM-assisted reviewing trial, calls for ECCV 2026 to follow — AntonObukhov1 · 2026-09-07
- Stolen API key uncovers Stratum, a Rust scanner sweeping 700,000 Docker layers a day for secrets — Ubunta · 2026-09-07
- Anthropic, Google, and OpenAI's $1 federal government contracts expire this month — LuizaJarovsky · 2026-09-07
- CodePen 2.0 sends editor input to its servers as you type, exposing unsaved secrets — maxim-fin · 2026-09-07
- The Jailbreak Argument Against LLM Values: Why Value Loading Isn't Solved — gleech · 2026-09-07