MiMo's GRS and GAR: scoring what makes an RL answer genuinely good, not just passing
tokenbender · x · 2026-09-22
Parts 7-8 of the thread cover the most interesting piece of the MiMo paper. Traditional RL treats pass=fail binary, yet two passing answers can differ wildly in quality. MiMo builds a better ruler, GRS: offline it reviews past rollouts to define what a genuinely good solution looks like, then uses that to improve grading. GAR (groupwise advantage redistribution) then splits credit among passing solutions that are merely okay, neat, or excellent.
Related event: MiMo v2.6 Redefines RL Rewarding with GRS and GAR(2 posts)→
More from Research
- Anthropic Is Setting Up a Biology Lab Where Claude Guides Robots Through Drug Experiments — The Decoder · 2026-09-22
- GAE: A Geometry-Native Autoencoder Cuts World Model FVD by 23.1% — yshan2u · 2026-09-22
- Over 1TB of China A-share Level-2 limit order book tick data hits Hugging Face — venvoo · 2026-09-22
- Reddit proposes measuring LLMs by cost per accepted task, not cost per token, after Grok 4.7 launch — Crescitaly · 2026-09-22
- Google's ScientistTwo solves 80.4% of 107 top-venue ML problems autonomously — thisdudelikesAI · 2026-09-22
- Tencent Hunyuan's WebCraftBench tests web apps like software, matching human preference 85.3% — TencentHunyuan · 2026-09-22