Shanghai AI Lab's AV-GRPO Uses Modality-Anchored RL to Beat LTX-2.3 at Joint Audio-Video Generation

Shanghai-AI-Laboratory · hf · 2026-09-25

Shanghai AI Laboratory introduces AV-GRPO, a modality-anchored online diffusion RL framework for joint audio-video generation, plus 5DAV, a decoupled, difficulty-controllable training dataset.

Existing joint audio-video models suffer from limited per-modality fidelity, weak text alignment, and poor cross-modal synchronization, and naive RL post-training runs into entangled heterogeneous rewards, costly joint two-tower optimization, and pair-dependent sync evaluation. AV-GRPO tackles these with three modules:

On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment, and cross-modal synchronization under both LoRA and full fine-tuning, with ablations confirming each design. Code and data are open-sourced.

Original post →

More from Multimodal

Multimodal channel →