Shanghai AI Lab's AV-GRPO Uses Modality-Anchored RL to Beat LTX-2.3 at Joint Audio-Video Generation
Shanghai-AI-Laboratory · hf · 2026-09-25
Shanghai AI Laboratory introduces AV-GRPO, a modality-anchored online diffusion RL framework for joint audio-video generation, plus 5DAV, a decoupled, difficulty-controllable training dataset.
Existing joint audio-video models suffer from limited per-modality fidelity, weak text alignment, and poor cross-modal synchronization, and naive RL post-training runs into entangled heterogeneous rewards, costly joint two-tower optimization, and pair-dependent sync evaluation. AV-GRPO tackles these with three modules:
- modality-anchored rollouts that disentangle learning signals and stabilize difficulty;
- trajectory-locked frozen-tower optimization to cut cost and reassign credit;
- adaptive objectives and perturbation strengths tailored to each modality's dynamics, converting coupled multimodal preference learning into unimodal subproblems.
On JavisBench and VABench, AV-GRPO outperforms LTX-2.3 in generation quality, semantic alignment, and cross-modal synchronization under both LoRA and full fine-tuning, with ablations confirming each design. Code and data are open-sourced.
More from Multimodal
- One-shot video generation gets shockingly good with a detailed shared prompt — TJLarkin23 · 2026-09-25
- Claude Opus 5.5 builds interactive 3D smelter model from drawings alone — EricBuess · 2026-09-25
- invideo launches Effects: build custom timeline effects by describing them in chat — Uncanny_Harry · 2026-09-25
- Creator lets Claude write its own response to p(doom) AI music videos — marcosalvi · 2026-09-25
- After generating 1,000+ songs, creator declares ElevenMusic 2.5 unmatched in AI music — kirbyman01 · 2026-09-25
- Sonder Editor adds reference management for MiniMax H3 video workflows in ComfyUI — SonderSaid · 2026-09-25