VA-Judger: human-preference reward model lifts joint video-audio generation to 62.3% human win rate
量子位 · wechat · 2026-09-21
Researchers from Fudan University and Shanghai Innovation Institute propose VA-Judger, a reward model for joint video-audio generation that judges the whole output instead of fragmented metrics (VideoAlign, AudioBox, CLAP, SynchFormer), which miss holistic issues and reward metric-gaming in RL post-training.
Approach
- Built VAPref-10K from 10k+ real clips (YouTube, Bilibili, film/TV): 9,000 prompts and 10k+ pairwise preferences, with annotators labeling both the winner and the deciding quality dimension.
- EasyPair (cross-model, clear gaps) for cold start; HardPair (same model, different seeds) for fine-grained alignment.
- Three-stage training: EasySFT for structured rubrics → HardSFT on hard pairs → Dimension-wise GRPO rewarding both correct picks and dimension-level scores.
Results
- Preference-prediction accuracy on VA-Judger-Bench: 68.43% vs 56.88% for the best single metric.
- Applied to LTX-2 post-training (8 candidates per prompt, 28 round-robin comparisons, frozen backbone + LoRA): ranks first on 11/13 JavisBench metrics; human three-way preference hits 62.30% vs 10.08% (LTX-2) and 27.63%.
Paper, code and demo are public.
More from Multimodal
- Turn any video into volumetric playback with exportable 360° orbits at custom tempo — LinusEkenstam · 2026-09-21
- Full Krea2 concept-art pipeline: four custom LoRAs, 5MP native frames, two-step upscale to 7000px — Dacrikka · 2026-09-21
- Phone video in, navigable 3D room out: K3-powered splat pipeline runs in the browser — willeastcott · 2026-09-21
- Viral Arabic AI trick: upload a photo and your name to get a magazine-cover portrait — aziz4ai · 2026-09-21
- MURAL METAMORPHOSIS: a fill-in prompt template that turns any subject into street-art murals — LudovicCreator · 2026-09-21
- Dreamina launches rebuilt web experience with film-grade workflows and unified canvas — Uncanny_Harry · 2026-09-21