V-Rubrics: rubric-based RL improves visual faithfulness in VLMs

liuziwei7 · x · 2026-08-29

A new paper, 'V-Rubrics: Visual Faithfulness via Rubric-Based Reinforcement Learning,' tackles vision-language models that answer fluently but wrongly: a single unsupported object, chart value, or inference step can invalidate an otherwise plausible response.

Problem: The authors frame this as a credit-assignment failure in multimodal post-training — scalar outcome rewards say whether an answer is acceptable, but not which visual facts are grounded or which reasoning steps are valid.

Method:

Results: Rubric-based GRPO beats both the shared SFT baseline and answer-only GRPO, with the largest gains on knowledge-oriented and visually grounded reasoning benchmarks — showing rubrics are a useful reward abstraction for visual post-training.

Original post →

More from Models

Models channel →