Policy gradient gets 4% vs 62% for cross-entropy on ImageNet, argues LLM post-training blame is misplaced

IanOsband · x · 2026-10-05

Ian Osband (DeepMind) argues that the community misdiagnoses LLM post-training failures by blaming classic RL difficulties—exploration, credit assignment, sampling noise. But policy gradient is exact in image classification, where none of those issues exist, and following it is still terrible: 4% on ImageNet versus 62% for cross-entropy. The problem lies with the policy gradient objective itself, not with RL's usual suspects.

Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→

Original post →

More from Research

Research channel →