NYU's RGPO uses adaptive rationale scaffolding to fight reward sparsity in RL for reasoning

newyorkuniversity · hf · 2026-10-07

NYU introduces Rationale-Guided Policy Optimization (RGPO), targeting reward sparsity in on-policy RL for LLM reasoning. Ground-truth rationales act as temporary scaffolds adapted to the model's current capability rather than fixed imitation targets; only higher-reward model-generated solutions are transferred back to the unguided setting, so off-policy data need not match the RL task format. RGPO consistently beats RLVR baselines in both text-only and vision-language reasoning, with ablations confirming adaptive guidance is key.

Original post →

More from Research

Research channel →