RLVR misses 'all minimal correct answers' problems; new credit assignment doubles finds
thoma_gu · x · 2026-10-09
T. Y. Tsui argues RLVR ignores a whole class of problems where the real requirement is finding ALL minimal sets of conditions that correctly produce an outcome — often misread as a preference for short or diverse outputs. Since existing RLVR scores each rollout independently, it can't tell minimal answers from redundant supersets or new alternatives from repeats. Derived from a chemistry reaction-space search task, the proposed credit assignment uses only the verifier's binary reward; in LLM post-training it finds about 2x as many minimal answers per problem within 64 samples as any baseline, including GRPO with a minimality oracle.
More from Research
- RWTH Aachen's ARROW unifies 3D reconstruction and point tracking from any RGB inputs, sets new SOTA — kwangmoo_yi · 2026-10-09
- HAIPS 2026 workshop lands at COLM tomorrow with top AI privacy researchers — tianshi_li · 2026-10-09
- Tighter Bayes error bounds via generalized Bhattacharyya and Chernoff means — FrnkNlsn · 2026-10-09
- PAMI: part-anchored motion lifts contact recall 14.5% in text-to-HOI generation — _akhaliq · 2026-10-09
- Study finds LLM novelty judges are unstable: scores swing wildly with evaluation design choices — _akhaliq · 2026-10-09
- Google's AMIE lands in The Lancet: 90% match with doctors' final diagnoses in real clinics — sundarpichai · 2026-10-09