Gaussian GRPO normalizes multimodal RL reward distributions, boosting OpenVLThinker v2

kaiwei_chang · x · 2026-10-06

Researchers unveiled Gaussian GRPO (G²RPO) at COLM 2026 to tackle a key challenge in post-training multimodal LLMs: reasoning-heavy and perception-heavy tasks yield very different reward distributions and learning dynamics. The method uses 1D optimal transport to force each task's advantage distribution to a standard normal, providing:

Combined with task-level response length and entropy shaping, it trains OpenVLThinker v2, which performs strongly on knowledge, math, chart, and document understanding benchmarks and outperforms closed-weight models over 100× larger on several evaluations, achieving SOTA open-source visual reasoning and perception among similar-size models.

Original post →

More from Multimodal

Multimodal channel →