DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL

DeepMind researcher Ian Osband released a paper, "Planning to Learn" (arXiv:2610.03667), arguing that the policy gradient commonly treated as the "RL loss" in LLM post-training has a fundamental flaw, and proposing a horizon loss implementable in one line of code that beats cross-entropy on MNIST and ImageNet.

Confirmed

Why it matters

2026-10-05 ~ 2026-10-05 · 5 related posts

Primary sources

1 near-duplicate retellings: IanOsband