VA-Judger: first reward model for joint video-audio generation from human preference

bdsqlsz · x · 2026-08-21

Researchers introduced VA-Judger, the first omni reward model for joint video-audio generation. The core problem: existing rewards combine separate metrics (audio quality, visual fidelity, sync), missing overall semantic and temporal coherence among prompt, video and audio — which invites reward hacking, producing high-scoring but incoherent outputs.

Key contributions:

Post-trained LTX-2 powered by VA-Judger generates content better aligned with human preferences.

Original post →

More from Multimodal

Multimodal channel →