Amazon & Duke study: supervising just 1 token per rollout boosts LLM reasoning
机器之心 · wechat · 2026-10-03
A new paper from Amazon and Duke University challenges the default assumption that post-training requires dense token-level supervision. Using On-Policy Distillation (OPD) as the testbed, they kept gradient signal at only 1 randomly chosen token per rollout (0.025% of thousands of tokens) — and student models still improved significantly on math reasoning.
Key findings
- Across 9 teacher-student configurations built on Qwen3 (small/large/equal sizes), single-token supervision consistently improved reasoning; results generalize to code reasoning, Llama models, and PPO-based RLVR.
- Selecting the token with the largest teacher-student disagreement: just 1-2 supervised tokens can match or beat full-signal training, with some students surpassing the teacher itself.
- Performance vs. supervision density is non-monotonic: it dips, then recovers and can exceed full-signal training at extreme sparsity, before collapsing when too sparse.
- Strongest students diverge more from teachers, implying gains don't come from imitation; 0.01% supervision updates only 10% of network parameters yet yields large gains.
- Positively-rewarded tokens work better when the student has strong representations; negatively-rewarded ones work better otherwise.
The authors draw an analogy to human learning: learners with prior knowledge need only sparse, well-placed corrections. Why a single token suffices remains open.
More from Research
- Policy Gradient for LLMs, Explained Visually: A From-Scratch REINFORCE Derivation — joecole · 2026-10-03
- The EinsteinTest: models given pre-breakthrough evidence rarely make the break themselves — CurieuxExplorer · 2026-10-03
- 4Director Lets You Direct AI Video by Placing Cameras and Objects in 3D — anand_bhattad · 2026-10-03
- JHU's GenCine turns a single image into an editable 3D scene for camera and object motion control in video generation — anand_bhattad · 2026-10-03
- Why SFT generalizes worse than RL: off-policy data, not the objective — a_karvonen · 2026-10-03
- Constant-Size Memory for Video World Models Proposed in New arXiv Paper — plsendfast · 2026-10-03