Tencent's RewardVerse uses rubric-guided optimization to fix video reward model drift

tencent · hf · 2026-09-24

Tencent researchers introduce RewardVerse, a rubric-based video reward framework that tackles scalar drift in reward models for video-generation RL. A dynamic rubric serves as an intermediate representation between query and scorer, while the two-stage RGPO algorithm jointly optimizes rubric generation and scorer alignment. Experiments on the 16-dimensional EvalVerse benchmark show state-of-the-art pointwise and pairwise evaluation with a stable, interpretable reward signal.

Original post →

More from Multimodal

Multimodal channel →