JustRL II adds token-level critic to GRPO for 128k long-CoT RL, lifting 2B model AIME25 from 61 to 81
zibuyu9 · x · 2026-09-08
- JustRL II released: In long-CoT (128k) RL, GRPO's group-mean baseline only gives a coarse response-level signal over tens of thousands of tokens.
- Method: Keep GRPO's group structure but add a critic for token-level credit assignment at similar simplicity; keeps improving where GRPO plateaus, taking a 2B model's AIME25 score from 61 to 81.
- Results: The same recipe powers the RL stage of MiniCPM5-2B, making it SOTA among models under 4B.
- Prior JustRL hit SOTA among 1.5B reasoning models with 2× less compute, stable over 4,000+ steps without multi-stage pipelines or dynamic schedules.
- Data and checkpoints are open; code lands this week.
More from Research
- VERA paper undercuts its own rhetoric: structured Jacobian-based IDM beats raw 14B video model for robot control — GeorgiaChal · 2026-09-08
- Researcher challenges scaling orthodoxy: benchmark trends aren't universal laws of perception and reasoning — GeorgiaChal · 2026-09-08
- If AI Helps Crack the Navier–Stokes Millennium Problem, It Would Signal a Scientific Revolution — kimmonismus · 2026-09-08
- New Preprint Unifies Integration and Access Accounts of Consciousness — anilkseth · 2026-09-08
- Mathematician argues papers should shift from proving conjectures to sharing insight — tak3sh8 · 2026-09-08
- Niantic's AutoCompass trains accurate visual localization from noisy GPS labels — ducha_aiki · 2026-09-08