VA-Judger: human-preference reward model lifts joint video-audio generation to 62.3% human win rate

量子位 · wechat · 2026-09-21

Researchers from Fudan University and Shanghai Innovation Institute propose VA-Judger, a reward model for joint video-audio generation that judges the whole output instead of fragmented metrics (VideoAlign, AudioBox, CLAP, SynchFormer), which miss holistic issues and reward metric-gaming in RL post-training.

Approach

Results

Paper, code and demo are public.

Original post →

More from Multimodal

Multimodal channel →