Research Unravels GRPO Training Collapse; ICML Highlights AI Scaling

Researchers have identified why the GRPO algorithm used in reasoning model training collapses—due to models losing confidence in their own outputs—and proposed a fix improving benchmarks by up to 45%. Meanwhile, ICML highlighted advancements in mathematical proof verification and long-horizon task planning.

2026-07-08 ~ 2026-07-08 · 2 related posts