Thinking Reward Model: rubric-first scoring sets open-source visual generation reward SOTA
Xuehai Bai · hf · 2026-09-30
KAIST introduces Think Before You Score, a paradigm where a reward model first builds case-adaptive rubrics before judging candidates. Their Thinking Reward Model (TRM) performs rubric-guided assessment and outputs fine-grained pointwise rewards. To fix score polarization from pairwise preference optimization, they propose Pairwise Dual-Group Relative Policy Optimization (PD-GRPO). TRM achieves state-of-the-art among open-source visual generation reward models, matching proprietary alternatives, and consistently improves diverse visual generation models when used as an RL reward signal.
More from Multimodal
- AutoRef open-sourced: harness optimization for agentic multi-reference image generation — NunyaBuzor · 2026-09-30
- Claude Opus 5.5 video imagines the moment Girl with a Pearl Earring was painted — justin_hart · 2026-09-30
- Reddit user shares AI-generated sci-fi short film Erebus-9: The Hiveborn Archives — himeonisama · 2026-09-30
- Teaching Claude Opus to Annotate Speech Timing for ElevenLabs Voiceovers Proves Tricky — RileyRalmuto · 2026-09-30
- Alibaba DAMO unveils WorldAttention for efficient interactive video world models — Alibaba-DAMO-Academy · 2026-09-30
- NUS's MaLiang-Harness exposes the Program-to-Visual gap in code-driven image/video generation — NationalUniversityofSingapore · 2026-09-30