V-Rubrics: 50k visual samples split into 353k checkable criteria fix multimodal RL credit assignment
jiqizhixin · x · 2026-10-04
NTU S-Lab, ASTAR, and UIUC present V-Rubrics, tackling credit assignment in multimodal RL: models can misread a chart number, hallucinate an object, and still land the correct final answer—result-scored RL rewards it anyway; conversely, a wrong answer doesn't mean every observation was wrong.
Key approach:
- Decomposes 50,248 visual samples into 352,938 individually checkable criteria;
- Visual facts, reasoning steps, and task requirements each participate separately in reward computation;
- Each criterion's score feeds GRPO reinforcement learning, reinforcing what was right and correcting what was not.
More from Multimodal
- First Song Made Entirely With AI: Suno, Midjourney, 116 Kling/Wan Clips — Afinetheorem · 2026-10-04
- A new Nano Banana Pro has been spotted; release timing unclear, likely not soon — koltregaskes · 2026-10-04
- Looped-DiT: 260M looped model beats 6.5x larger text-to-image rival with 4.9x less compute — Apprehensive_Sky892 · 2026-10-04
- ComfyUI nodes add reference-image and phrase-level attention control to Qwen Image 2.1 — Capitan01R- · 2026-10-04
- Eleven v4 Delivers First Zero-Edit Voice Output in Blogger's Test, 9 Free Days Left — FellMentKE · 2026-10-04
- AdScanVideo Turns Any Uploaded Video Into a Timestamped, Structured Analysis Report — aftahi_ai · 2026-10-04