NJU Proposes CSCR: Fixing Credit Assignment in Long-CoT RL

NJU · hf · 2026-08-03

A research team from Nanjing University points out a flaw in current Reinforcement Learning with Verifiable Rewards (RLVR, e.g., GRPO) when handling long Chain-of-Thought (Long-CoT) reasoning: they uniformly broadcast response-level rewards across all tokens, ignoring their unequal actual contributions to the final outcome.

The team found that likelihood shifts induced by On-policy self-distillation (OPSD) primarily concentrate on highly substitutable surface-form tokens rather than those carrying actual reasoning logic. To address this, they propose Counterfactual Sensitivity Credit Reallocation (CSCR):

Original post →

More from Research

Research channel →