Gaussian GRPO normalizes multimodal RL reward distributions, boosting OpenVLThinker v2
kaiwei_chang · x · 2026-10-06
Researchers unveiled Gaussian GRPO (G²RPO) at COLM 2026 to tackle a key challenge in post-training multimodal LLMs: reasoning-heavy and perception-heavy tasks yield very different reward distributions and learning dynamics. The method uses 1D optimal transport to force each task's advantage distribution to a standard normal, providing:
- intrinsic robustness to outliers
- symmetric updates for positive and negative rewards
- uniform variance across diverse tasks
Combined with task-level response length and entropy shaping, it trains OpenVLThinker v2, which performs strongly on knowledge, math, chart, and document understanding benchmarks and outperforms closed-weight models over 100× larger on several evaluations, achieving SOTA open-source visual reasoning and perception among similar-size models.
More from Multimodal
- AI-generated video is so funny netizens say the compute was worth it — hexiang · 2026-10-06
- Midjourney + Threejs + Seedance workflow for multi-angle AI cinematography — Ror_Fly · 2026-10-06
- Suno turns pictures into songs — and it even recognized the Gram-Schmidt meme — anderssandberg · 2026-10-06
- Apple researcher lands two NeurIPS 2026 papers on normalizing flows, including Normalizing Trajectory Models — thoma_gu · 2026-10-06
- SemanTok: 49M autoregressive video model beats SOTA rivals 47x its size — CSProfKGD · 2026-10-06
- Magnific teases October 8 launch, simple prompts already yield impressive results — aziz4ai · 2026-10-06