Ian Osband's Planning to Learn: One-Line Horizon Loss Tops Cross-Entropy on MNIST and ImageNet

IanOsband · x · 2026-10-05

DeepMind researcher Ian Osband published Planning to Learn (arXiv:2610.03667), reframing the relation between policy gradient and classification loss:

Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→

Original post →

More from Research

Research channel →