Ian Osband: policy gradient hits 4% vs 62% cross-entropy on ImageNet — RL loss isn't the problem
IanOsband · x · 2026-10-05
DeepMind researcher Ian Osband argues that LLM post-training treats policy gradient as "the RL loss" and blames failures on exploration, credit assignment and sampling noise — but even in image classification, where none of those exist, strictly following the exact policy gradient is terrible: 4% vs 62% for cross-entropy on ImageNet.
He illustrates with two examples: A) correct label at p=0.9, B) at p=1e-6. A tiny policy-gradient step sees no point improving B (success rate too small), yet over many steps investing in B beats A.
He frames cross-entropy as "patient accuracy": the total error an example would pay if training never stopped, while policy gradient is the zero-horizon limit. Cutting the total off at the remaining budget yields a "horizon loss" — implementable in one line of code.
Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→
More from Research
- PCA vs EGA on LLM embeddings: EGA recovers 6-dim structure 97.6–100%, PCA nearly 0% — GolinoHudson · 2026-10-06
- Stanford studies find AI benchmarks may not measure what they claim, with billions riding on scores — StanfordHAI · 2026-10-06
- Reverse engineering Neuralink's new decoder architecture from a sparse blog post — melnykowycz · 2026-10-06
- Harvard-MIT paper: "plan ahead" prompts make agents play worse; interface design matters more — dair_ai · 2026-10-06
- First Human–AI Interaction Conference HAIC 2027 announced for June in Washington, DC — jasonwuishere · 2026-10-06
- Gym-Anything gets oral presentation at COLM 2026 Lifelong Agents Workshop — wellecks · 2026-10-06