MiMo's GRS and GAR: scoring what makes an RL answer genuinely good, not just passing

tokenbender · x · 2026-09-22

Parts 7-8 of the thread cover the most interesting piece of the MiMo paper. Traditional RL treats pass=fail binary, yet two passing answers can differ wildly in quality. MiMo builds a better ruler, GRS: offline it reviews past rollouts to define what a genuinely good solution looks like, then uses that to improve grading. GAR (groupwise advantage redistribution) then splits credit among passing solutions that are merely okay, neat, or excellent.

Related event: MiMo v2.6 Redefines RL Rewarding with GRS and GAR(2 posts)→

Original post →

More from Research

Research channel →