Ian Osband's Planning to Learn: One-Line Horizon Loss Tops Cross-Entropy on MNIST and ImageNet
IanOsband · x · 2026-10-05
DeepMind researcher Ian Osband published Planning to Learn (arXiv:2610.03667), reframing the relation between policy gradient and classification loss:
- Problem: exact policy gradient in classification is smooth and noise-free yet still loses to cross-entropy—because the exact gradient is myopic, valuing updates only by immediate gains while each update also sets where the next one starts.
- Insight: cross-entropy is 'patient accuracy' (the total error an example would pay if its log-odds rose at unit speed forever); exact policy gradient is its zero-horizon limit.
- Method: truncating that total at the remaining learning budget yields the horizon loss—a one-line change that moves from cross-entropy toward exact policy gradient as training runs out.
- Results: provably escapes the traps of both endpoints in an allocation model; improves top-1 accuracy over cross-entropy on MNIST and ImageNet (ResNet-50/101, ViT-S/16) at flat learning rates, with gains growing under label noise.
Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→
More from Research
- Researchers warn: PCA misestimates personality structure in LLM embeddings, EGA works better — GolinoHudson · 2026-10-06
- Stanford studies find AI benchmarks may not measure what they claim, with billions riding on scores — StanfordHAI · 2026-10-06
- Reverse engineering Neuralink's new decoder architecture from a sparse blog post — melnykowycz · 2026-10-06
- Harvard-MIT paper: "plan ahead" prompts make agents play worse; interface design matters more — dair_ai · 2026-10-06
- First Human–AI Interaction Conference HAIC 2027 announced for June in Washington, DC — jasonwuishere · 2026-10-06
- Gym-Anything gets oral presentation at COLM 2026 Lifelong Agents Workshop — wellecks · 2026-10-06