NYU's RGPO uses adaptive rationale scaffolding to fight reward sparsity in RL for reasoning
newyorkuniversity · hf · 2026-10-07
NYU introduces Rationale-Guided Policy Optimization (RGPO), targeting reward sparsity in on-policy RL for LLM reasoning. Ground-truth rationales act as temporary scaffolds adapted to the model's current capability rather than fixed imitation targets; only higher-reward model-generated solutions are transferred back to the unguided setting, so off-policy data need not match the RL task format. RGPO consistently beats RLVR baselines in both text-only and vision-language reasoning, with ablations confirming adaptive guidance is key.
More from Research
- Naive baseline beats AI model in ~95% of cases, exposing flaws in biology benchmarks — bravo_abad · 2026-10-07
- The scaling debate hinges on two readings of "predictably better with scale" — Diyi_Yang · 2026-10-07
- HarnessTester finds 100+ real bugs in LLM agent harnesses like OpenClaw — LingmingZhang · 2026-10-07
- AWS paper: structured agent communication (AECP) lifts multi-agent coding pass rate 28.2% — dair_ai · 2026-10-07
- Does RSI help with data? Frontier model progress may hinge on data, not compute — maxsloef · 2026-10-07
- Bug Hunt Bench: Mistral Large 4 fixes only 15/105 planted bugs, trails Qwen and Kimi — PawelHuryn · 2026-10-07