RecCAR closes reciprocal cross-attention gap in joint video diffusion models
barilan · hf · 2026-09-24
- The study finds joint multimodal diffusion transformers are asymmetric in cross-modal correspondence: companion modalities (3D body motion, audio) strongly track video, but their reciprocal constraints on video remain weak — the "reciprocal correspondence gap."
- RecCAR (Reciprocal Cross-modal Attention Regularization) is a KL regularizer aligning the weaker modality-to-video correspondence toward the stable video-to-modality distribution.
- On joint video-motion and video-audio generation, Human Anatomy score rises from 0.69 to 0.75 and audio-video desynchronization drops from 0.804 to 0.752, with overall generation improved.
More from Multimodal
- Qwen Image 2.1 falls back to image editing when given reference images — extra2AB · 2026-09-24
- Kling 4.0 leak: 120-second videos, Omni Reference with 15 elements, up to 4K images — koltregaskes · 2026-09-24
- Limestone claims Claude Opus 5.5 one-shotted its entire launch video — alex_verem · 2026-09-24
- Using Ling-3.0-flash-VL to analyze 8 ultrasound reports and solve a pregnancy puzzle — alifcoder · 2026-09-24
- Claude synthesizes a full band from math, noise and its own voice — willcb · 2026-09-24
- Qwen Image 2.1 secretly adds alpha channels, inflating file size by 15-20% — Calm_Mix_3776 · 2026-09-24