NVIDIA's HDL Cuts RLVR Token Cost 2.5x by Localizing Where Models Change Their Mind
nvidia · hf · 2026-09-30
NVIDIA proposes Hindsight-Divergence Localization (HDL) for RL with verifiable rewards: hindsight-induced shifts in token log-likelihoods identify branch points where the policy reconsiders earlier choices, and training groups are filled with suffix-only continuations from those positions. Across three models on math, code, and agent tasks, HDL reduces generated tokens up to 2.5x and speeds rollout 1.8x versus GRPO, while improving performance—up to 12.5 points on agent tasks.
More from Research
- 16 Mathematicians Publish Leiden Declaration on AI's Role in Mathematics Research — burny_tech · 2026-09-30
- Repeated Solutions Make Reasoning Fragile After Instruction Tuning, Paper Finds — burny_tech · 2026-09-30
- AI is the discipline of building minds, and mathematics is just getting started — burny_tech · 2026-09-30
- A new auto-research loop that bootstraps the shape of the best possible result — burny_tech · 2026-09-30
- EasyPPO: just fix the critic — stable PPO for LLM post-training with zero training collapse — teortaxesTex · 2026-09-30
- Looped MoE Tuning: 2x Experts, 0.5x Looped Layers, 2x Loops, Attention Untied — burny_tech · 2026-09-30