Tencent's RewardVerse uses rubric-guided optimization to fix video reward model drift
tencent · hf · 2026-09-24
Tencent researchers introduce RewardVerse, a rubric-based video reward framework that tackles scalar drift in reward models for video-generation RL. A dynamic rubric serves as an intermediate representation between query and scorer, while the two-stage RGPO algorithm jointly optimizes rubric generation and scorer alignment. Experiments on the 16-dimensional EvalVerse benchmark show state-of-the-art pointwise and pairwise evaluation with a stable, interpretable reward signal.
More from Multimodal
- Redditor Shares Qwen 2.1 Character LoKr Training Recipe: Multi-Resolution Is Key — Any_Tea_3499 · 2026-09-24
- Open-Source MCP Server Connects Claude/Cursor to ComfyUI with 50+ Tools — Less_Actuary_9441 · 2026-09-24
- From a rough 15-second demo to a cinematic 30-second SpaceX concept with Pexo — nikola_mr64990 · 2026-09-24
- Blender MCP hits 29k stars as Opus 5.5 builds cities on medium effort — sidahuj · 2026-09-24
- Fish Audio launches Drama 3 preview, billing it as the most controllable TTS model ever — bdsqlsz · 2026-09-24
- Redditor Builds Five-Scene Guinea Pig Documentary with Google Veo and Phonetic Sync — No_Ruin_3716 · 2026-09-24