VA-Judger: first reward model for joint video-audio generation from human preference
bdsqlsz · x · 2026-08-21
Researchers introduced VA-Judger, the first omni reward model for joint video-audio generation. The core problem: existing rewards combine separate metrics (audio quality, visual fidelity, sync), missing overall semantic and temporal coherence among prompt, video and audio — which invites reward hacking, producing high-scoring but incoherent outputs.
Key contributions:
- VAPref-10K: a human-preference dataset with 9K prompts and 10.3K fine-grained paired comparisons from open-source generation models
- VA-Judger-Bench: a benchmark with in-domain and out-of-domain comparisons to test whether reward models truly align with human preferences
- VA-Judger uses chain-of-thought reward modeling: first learning from pairs with clear quality gaps for structured outputs and coarse preference, then distilling finer judgments
Post-trained LTX-2 powered by VA-Judger generates content better aligned with human preferences.
More from Multimodal
- ByteDance research improves image generation with structured prompts — jiqizhixin · 2026-08-21
- Tutorial: Character creation with elements in Adobe Firefly — LudovicCreator · 2026-08-21
- Fixing LTX 2.5 smearing: Ported custom nodes and workflow — SillyLilithh · 2026-08-21
- Senior Art Director Praises Grok Imagine 2: Saves Time by Converting Screenshots to Game Assets — hexiang · 2026-08-21
- Fairground AI Creator TV channel launches on Amazon Prime Video — dolma33 · 2026-08-21
- Local Video Generation with MiniMax H3 Enables Seamless Extension — cocktailpeanut · 2026-08-21