DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL
DeepMind researcher Ian Osband released a paper, "Planning to Learn" (arXiv:2610.03667), arguing that the policy gradient commonly treated as the "RL loss" in LLM post-training has a fundamental flaw, and proposing a horizon loss implementable in one line of code that beats cross-entropy on MNIST and ImageNet.
Confirmed
- Osband pointed out that when post-training fails, people habitually blame classic RL difficulties—exploration, credit assignment, sampling noise—but in image classification, where none of these issues exist, the exact (smooth, noiseless) policy gradient still lags far behind: on ImageNet, policy gradient reaches only about 4% versus 62% for cross-entropy.
- He interprets cross-entropy as "patient accuracy": the total error a sample would eventually incur if trained indefinitely, assuming log-odds keep rising at unit speed; policy gradient is the zero-horizon limit of that total.
- Horizon loss is defined by truncating that total at the remaining training budget. The paper reports that this loss beats cross-entropy on MNIST and ImageNet.
- The work offers a unified view of classification losses and reinforcement learning, suggesting RL and supervised learning can be understood within the same framework.
Why it matters
- If the "myopia" of policy gradients is indeed a fundamental flaw, the core loss functions underlying LLM post-training (RLHF/RLVR, etc.) may not be optimal, and horizon loss may offer a more direct path to improvement.
- The paper unifies supervised learning and RL losses in one mathematical framework, with potential implications for both theoretical understanding and engineering practice.
2026-10-05 ~ 2026-10-05 · 5 related posts
Primary sources
- Policy gradient gets 4% vs 62% for cross-entropy on ImageNet, argues LLM post-training blame is misplaced — IanOsband · 2026-10-05
- Ian Osband: policy gradient hits 4% vs 62% cross-entropy on ImageNet — RL loss isn't the problem — IanOsband · 2026-10-05
- Ian Osband's 'horizon loss': truncating cross-entropy by remaining training budget in one line — IanOsband · 2026-10-05
- Ian Osband's 'Planning to Learn': One-Line Horizon Loss Beats Both Policy Gradient and Cross-Entropy — IanOsband · 2026-10-05
1 near-duplicate retellings: IanOsband