Ian Osband: policy gradient hits 4% vs 62% cross-entropy on ImageNet — RL loss isn't the problem

IanOsband · x · 2026-10-05

DeepMind researcher Ian Osband argues that LLM post-training treats policy gradient as "the RL loss" and blames failures on exploration, credit assignment and sampling noise — but even in image classification, where none of those exist, strictly following the exact policy gradient is terrible: 4% vs 62% for cross-entropy on ImageNet.

He illustrates with two examples: A) correct label at p=0.9, B) at p=1e-6. A tiny policy-gradient step sees no point improving B (success rate too small), yet over many steps investing in B beats A.

He frames cross-entropy as "patient accuracy": the total error an example would pay if training never stopped, while policy gradient is the zero-horizon limit. Cutting the total off at the remaining budget yields a "horizon loss" — implementable in one line of code.

Related event: DeepMind's Ian Osband Proposes Horizon Loss Unifying Classification and RL(5 posts)→

Original post →

More from Research

Research channel →