Ian Osband's 'Planning to Learn': One-Line Horizon Loss Beats Both Policy Gradient and Cross-Entropy

IanOsband · x · 2026-10-05

Ian Osband introduces the core idea of his paper Planning to Learn: exact policy gradient is myopic, valuing updates only by immediate gains, while cross-entropy is 'patient accuracy'—the total error an example would pay if its log-odds rose at unit speed forever, i.e. the zero-horizon limit. Truncating that total at the remaining learning budget yields the horizon loss, a one-line change applicable to cross-entropy too. Key insight: 'the value of an update is not necessarily just the gradient of a loss.'

Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→

Original post →

More from Research

Research channel →