Ian Osband's 'horizon loss': truncating cross-entropy by remaining training budget in one line

IanOsband · x · 2026-10-05

Follow-up to Ian Osband's thread: he defines cross-entropy as "patient accuracy" — the total error an example would pay if training on it never stopped — while policy gradient is the zero-horizon limit of that total. Truncating the total by the remaining training budget yields the "horizon loss", implementable in one line of code.

This closes his argument for why policy gradient scores 4% vs 62% cross-entropy on ImageNet: it abandons the long-horizon view, and horizon loss offers a one-line compromise.

Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→

Original post →

More from Research

Research channel →