NJU Proposes CSCR: Fixing Credit Assignment in Long-CoT RL
NJU · hf · 2026-08-03
A research team from Nanjing University points out a flaw in current Reinforcement Learning with Verifiable Rewards (RLVR, e.g., GRPO) when handling long Chain-of-Thought (Long-CoT) reasoning: they uniformly broadcast response-level rewards across all tokens, ignoring their unequal actual contributions to the final outcome.
The team found that likelihood shifts induced by On-policy self-distillation (OPSD) primarily concentrate on highly substitutable surface-form tokens rather than those carrying actual reasoning logic. To address this, they propose Counterfactual Sensitivity Credit Reallocation (CSCR):
- Core Mechanism: As an extension of GRPO, it reduces credit for highly sensitive tokens and renormalizes token-level advantages to preserve the original credit budget and verifier-determined direction.
- Results: On long-CoT mathematical reasoning benchmarks, CSCR consistently outperforms the GRPO baseline with the same number of policy updates. Ablations confirm that moderate downweighting is most effective, while stronger modulation destabilizes optimization.
More from Research
- Applications Open for MLSS 2027 Okinawa with Top AI Researchers Announced — FrnkNlsn · 2026-08-03
- New Benchmark Just Dropped — nptacek · 2026-08-03
- Two Teams Solve Same Quantum Problem with GPT-5.6, Blurring 'Independent Discovery' — The Decoder · 2026-08-03
- Gradian: Open-Source Tool to Pinpoint Poison Data Causing LLM Fine-Tuning Failures — vylara-ai · 2026-08-03
- Explained: How AI Companies Achieve 10x Faster Video Model Inference — haremlifegame · 2026-08-03
- Echoverse: Training Agents in Deep Environments Boosts 9B Model Success Rate to 67% — burny_tech · 2026-08-03