HKUST's CorrGRPO Normalizes Reward Correlations to Fix Multi-Reward GRPO Training
HKUST · hf · 2026-10-02
HKUST proposes CorrGRPO, an improvement to GRPO for multi-reward RL training. Standard GRPO sums rewards and normalizes by within-group standard deviation, so large-scale correlated rewards can dominate the normalization and suppress smaller-scale reward signals.
CorrGRPO normalizes pairwise covariances into Pearson correlation coefficients, keeping the centered total reward unchanged while letting advantage magnitudes adapt to reward correlations. It outperforms GRPO and variants on code generation, tool calling, and agent security tasks across 0.5B–8B models. Code is open-sourced on GitHub.
More from Research
- Sasha Rush heads to COLM 2025, inviting chats on TTT, proofs and biased RL — srush_nlp · 2026-10-02
- CISPA Offers Imprecise Probabilistic ML Course Again, Free and Online — krikamol · 2026-10-02
- SoftServe preprint brings scalable quasi-Newton optimization to deep learning, beating Adam, Muon and SOAP on ill-conditioned tasks — dianarycai · 2026-10-02
- SoftServe's updates build on quadratic matrix equations, extending the authors' Batch-and-Match BBVI work — dianarycai · 2026-10-02
- DyRAD: radar novel view synthesis renders full range-azimuth-Doppler tensors for dynamic driving scenes — orlitany · 2026-10-02
- EMBL-EBI's saezlab open-sources Karenina, a framework for multi-dimensional biomedical AI evaluation — anshulkundaje · 2026-10-02