Amazon & Duke study: supervising just 1 token per rollout boosts LLM reasoning

机器之心 · wechat · 2026-10-03

A new paper from Amazon and Duke University challenges the default assumption that post-training requires dense token-level supervision. Using On-Policy Distillation (OPD) as the testbed, they kept gradient signal at only 1 randomly chosen token per rollout (0.025% of thousands of tokens) — and student models still improved significantly on math reasoning.

Key findings

The authors draw an analogy to human learning: learners with prior knowledge need only sparse, well-placed corrections. Why a single token suffices remains open.

Original post →

More from Research

Research channel →