Ian Osband's 'Planning to Learn': One-Line Horizon Loss Beats Both Policy Gradient and Cross-Entropy
IanOsband · x · 2026-10-05
Ian Osband introduces the core idea of his paper Planning to Learn: exact policy gradient is myopic, valuing updates only by immediate gains, while cross-entropy is 'patient accuracy'—the total error an example would pay if its log-odds rose at unit speed forever, i.e. the zero-horizon limit. Truncating that total at the remaining learning budget yields the horizon loss, a one-line change applicable to cross-entropy too. Key insight: 'the value of an update is not necessarily just the gradient of a loss.'
Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→
More from Research
- Frank Nielsen's free textbook: An Elementary Introduction to Information Geometry — FrnkNlsn · 2026-10-05
- Microsoft's ThinkingBox: grade agents by database changes, not their words — SergioPaniego · 2026-10-05
- Quanta: Is AI the End of Math As We Know It? — littmath · 2026-10-05
- First-of-Kind RCT: GPT-4o Respiratory Chatbot Beats Web Search for Layperson Diagnosis — EricTopol · 2026-10-05
- Local sparsity enables unsupervised LLM safety detection, new NeurIPS paper shows — breadli428 · 2026-10-05
- Eric Topol highlights major RAS pancreatic cancer advance, MYC next in line — EricTopol · 2026-10-05