Tencent's Adaptive Reward Routing Dynamically Balances Multi-Reward RL for Joint Audio-Video Diffusion

tencent · hf · 2026-10-02

Tencent proposes Adaptive Reward Routing for forward-process RL (DiffusionNFT) of joint audio-video diffusion models, dynamically adapting where reward updates act and how competing rewards combine. Cross-Modal Influence-Guided Routing uses bidirectional cross-attention responses as a proxy to reweight token-level losses and scale gradients across cross-modal layers; Preference-Preserving Modality-Aware Reweighting keeps predefined weights as priors with residual corrections from branch-level reward-gradient interactions. Experiments show consistent gains in modality quality, semantic consistency, and audio-video synchronization over strong RL baselines.

Original post →

More from Multimodal

Multimodal channel →