NVIDIA's HDL Cuts RLVR Token Cost 2.5x by Localizing Where Models Change Their Mind

nvidia · hf · 2026-09-30

NVIDIA proposes Hindsight-Divergence Localization (HDL) for RL with verifiable rewards: hindsight-induced shifts in token log-likelihoods identify branch points where the policy reconsiders earlier choices, and training groups are filled with suffix-only continuations from those positions. Across three models on math, code, and agent tasks, HDL reduces generated tokens up to 2.5x and speeds rollout 1.8x versus GRPO, while improving performance—up to 12.5 points on agent tasks.

Original post →

More from Research

Research channel →