Thinking Reward Model: rubric-first scoring sets open-source visual generation reward SOTA

Xuehai Bai · hf · 2026-09-30

KAIST introduces Think Before You Score, a paradigm where a reward model first builds case-adaptive rubrics before judging candidates. Their Thinking Reward Model (TRM) performs rubric-guided assessment and outputs fine-grained pointwise rewards. To fix score polarization from pairwise preference optimization, they propose Pairwise Dual-Group Relative Policy Optimization (PD-GRPO). TRM achieves state-of-the-art among open-source visual generation reward models, matching proprietary alternatives, and consistently improves diverse visual generation models when used as an RL reward signal.

Original post →

More from Multimodal

Multimodal channel →