HKUST's CorrGRPO Normalizes Reward Correlations to Fix Multi-Reward GRPO Training

HKUST · hf · 2026-10-02

HKUST proposes CorrGRPO, an improvement to GRPO for multi-reward RL training. Standard GRPO sums rewards and normalizes by within-group standard deviation, so large-scale correlated rewards can dominate the normalization and suppress smaller-scale reward signals.

CorrGRPO normalizes pairwise covariances into Pearson correlation coefficients, keeping the centered total reward unchanged while letting advantage magnitudes adapt to reward correlations. It outperforms GRPO and variants on code generation, tool calling, and agent security tasks across 0.5B–8B models. Code is open-sourced on GitHub.

Original post →

More from Research

Research channel →