Adaptive Reward Routing improves joint audio-video diffusion via forward-process RL
_akhaliq · x · 2026-10-03
This paper proposes Adaptive Reward Routing for dynamically adapting update locations and reward coordination during forward-process RL (DiffusionNFT) of joint audio-video diffusion models. Two components: (1) Cross-Modal Influence-Guided Routing uses bidirectional cross-attention responses as a proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without extra model interventions; (2) Preference-Preserving Modality-Aware Reweighting keeps predefined weights as preference priors and applies branch-specific reward-gradient interactions as residual corrections after warm-up, preventing dominant rewards from suppressing weak-but-essential objectives. Experiments show consistent gains in modality quality, semantic consistency, and audio-video sync over strong RL baselines, validated by ablations and mechanism analyses.
Related event: Tencent Proposes Adaptive Reward Routing for Joint Audio-Video Diffusion RL(2 posts)→
More from Multimodal
- Bona Film's 100-minute AI-assisted feature 'Sanxingdui: Future Past' gets theatrical release Oct 23 — lmoroney · 2026-10-03
- AI Director Junie Lau Featured in Forbes; Goldfish Mayhem Premieres with Wonder Studios — JunieLauX · 2026-10-03
- fal's H3 Max Reference-to-Video adds first, middle, and last frame control — OdinLovis · 2026-10-03
- Developer Runs Blender Inside ChatGPT Dots to Auto-Model, Rig and Animate a Dance Video — OpenAIDevs · 2026-10-03
- Tavus's Griffin starts talking before generation finishes, face reacts in 0.43s — TejasKumar_ · 2026-10-03
- Patching ComfyUI-MPS-INT8 to fix MiniMax H3 VAE crashes on Mac — Kris_RenderBob · 2026-10-03