Adaptive Reward Routing improves joint audio-video diffusion via forward-process RL

_akhaliq · x · 2026-10-03

This paper proposes Adaptive Reward Routing for dynamically adapting update locations and reward coordination during forward-process RL (DiffusionNFT) of joint audio-video diffusion models. Two components: (1) Cross-Modal Influence-Guided Routing uses bidirectional cross-attention responses as a proxy for evolving cross-modal influence, dynamically reweighting token-aware losses and scaling gradients across cross-modal layers without extra model interventions; (2) Preference-Preserving Modality-Aware Reweighting keeps predefined weights as preference priors and applies branch-specific reward-gradient interactions as residual corrections after warm-up, preventing dominant rewards from suppressing weak-but-essential objectives. Experiments show consistent gains in modality quality, semantic consistency, and audio-video sync over strong RL baselines, validated by ablations and mechanism analyses.

Related event: Tencent Proposes Adaptive Reward Routing for Joint Audio-Video Diffusion RL(2 posts)→

Original post →

More from Multimodal

Multimodal channel →