RLVR misses 'all minimal correct answers' problems; new credit assignment doubles finds

thoma_gu · x · 2026-10-09

T. Y. Tsui argues RLVR ignores a whole class of problems where the real requirement is finding ALL minimal sets of conditions that correctly produce an outcome — often misread as a preference for short or diverse outputs. Since existing RLVR scores each rollout independently, it can't tell minimal answers from redundant supersets or new alternatives from repeats. Derived from a chemistry reaction-space search task, the proposed credit assignment uses only the verifier's binary reward; in LLM post-training it finds about 2x as many minimal answers per problem within 64 samples as any baseline, including GRPO with a minimality oracle.

Original post →

More from Research

Research channel →