Policy gradient gets 4% vs 62% for cross-entropy on ImageNet, argues LLM post-training blame is misplaced
IanOsband · x · 2026-10-05
Ian Osband (DeepMind) argues that the community misdiagnoses LLM post-training failures by blaming classic RL difficulties—exploration, credit assignment, sampling noise. But policy gradient is exact in image classification, where none of those issues exist, and following it is still terrible: 4% on ImageNet versus 62% for cross-entropy. The problem lies with the policy gradient objective itself, not with RL's usual suspects.
Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→
More from Research
- Study: DeepSeek's mHC Uses Only Two of Four Residual Streams, Deep Mixing Goes Near-Identity — jiqizhixin · 2026-10-05
- Frank Nielsen's free textbook: An Elementary Introduction to Information Geometry — FrnkNlsn · 2026-10-05
- Microsoft's ThinkingBox: grade agents by database changes, not their words — SergioPaniego · 2026-10-05
- Quanta: Is AI the End of Math As We Know It? — littmath · 2026-10-05
- First-of-Kind RCT: GPT-4o Respiratory Chatbot Beats Web Search for Layperson Diagnosis — EricTopol · 2026-10-05
- Local sparsity enables unsupervised LLM safety detection, new NeurIPS paper shows — breadli428 · 2026-10-05