Goal Misgeneralization in Deep RL: Agents Keep Their Skills but Pursue the Wrong Goal

gleech · x · 2026-09-07

gleech cites the ICML 2022 paper "Goal Misgeneralization in Deep Reinforcement Learning" (Langosco, Koch, Sharkey, Pfau, Orseau, Krueger) as empirical support for the thread: goal misgeneralization is an out-of-distribution failure where an agent retains its capabilities yet pursues the wrong goal — e.g., competently avoiding obstacles while navigating to the wrong place. Unlike prior work on capability generalization failures, the paper formalizes the distinction between capability and goal generalization, provides the first empirical demonstrations, and partially characterizes its causes.

Original post →

More from Safety

Safety channel →