Ian Osband's 'horizon loss': truncating cross-entropy by remaining training budget in one line
IanOsband · x · 2026-10-05
Follow-up to Ian Osband's thread: he defines cross-entropy as "patient accuracy" — the total error an example would pay if training on it never stopped — while policy gradient is the zero-horizon limit of that total. Truncating the total by the remaining training budget yields the "horizon loss", implementable in one line of code.
This closes his argument for why policy gradient scores 4% vs 62% cross-entropy on ImageNet: it abandons the long-horizon view, and horizon loss offers a one-line compromise.
Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→
More from Research
- PCA vs EGA on LLM embeddings: EGA recovers 6-dim structure 97.6–100%, PCA nearly 0% — GolinoHudson · 2026-10-06
- Stanford studies find AI benchmarks may not measure what they claim, with billions riding on scores — StanfordHAI · 2026-10-06
- Reverse engineering Neuralink's new decoder architecture from a sparse blog post — melnykowycz · 2026-10-06
- Harvard-MIT paper: "plan ahead" prompts make agents play worse; interface design matters more — dair_ai · 2026-10-06
- First Human–AI Interaction Conference HAIC 2027 announced for June in Washington, DC — jasonwuishere · 2026-10-06
- Gym-Anything gets oral presentation at COLM 2026 Lifelong Agents Workshop — wellecks · 2026-10-06