Progressive Point Matching: Dense Per-Token Credit for Long-Horizon LLM RL, Beating GRPO Exponentially
kvfrans · x · 2026-09-09
- Problem: Standard LLM RL uses sparse outcome rewards (0/1 per trajectory), which grow exponentially inefficient as tasks scale to millions or billions of tokens — theoretically, policy gradient signal-to-noise degrades exponentially with task horizon. A trajectory completing dozens of subtasks but failing the last step earns the same reward as one that made no progress.
- Why existing fixes fall short: Learned value functions, process rewards, and self-distillation introduce asymptotic bias — process rewards can reward logically correct statements irrelevant to task success.
- The method: Progressive Point Matching (PPM) sits between RL and imitation learning, providing a simple, asymptotically unbiased way to assign dense per-token credit over long trajectories.
- Result: The authors report exponentially improved training efficiency over standard GRPO on long-horizon tasks.
- Written by Preston Fu and colleagues (including Kevin Frans); full details in the accompanying blog post.
More from Research
- Lightwheel open-sources 100,000 hours of egocentric video for robot training — chris_j_paxton · 2026-09-09
- Spilled.ink launches crowd-voted tracker for the OpenAI Navier–Stokes proof debate — NathanpmYoung · 2026-09-09
- DNA sequence models hit SOTA but wet-lab validation remains the bottleneck, experts say — anshulkundaje · 2026-09-09
- Researcher says he predicted Navier-Stokes would fall first, but not amid bitter controversy — geoffreyirving · 2026-09-09
- Stanford researcher: mathematicians expected a Navier-Stokes counterexample was the likely outcome — RishiBommasani · 2026-09-09
- OpenAI's Internal Model Cracks Open Math Problems Where Humanity Scored Effectively Zero — stevenheidel · 2026-09-09