Tencent Hunyuan's training-free MC-Sparse attention speeds DiT denoising up to 2.3x
Tencent-Hunyuan · hf · 2026-10-09
Tencent Hunyuan released MC-Sparse (Meta-Cached Sparse Attention), a training-free sparse attention framework for accelerating diffusion transformers in long-sequence generation like video and high-res 3D assets.
- Controlled oracle comparisons trace quality degradation at high sparsity to three sources: token grouping constraints, inaccurate interaction selection, and attention contributions lost from discarded tokens.
- MC-Sparse selects individual KV tokens, groups similar queries into tile-aligned bins for GPU efficiency, and caches query groups, exact-probability-selected KV indices, and dense-sparse residuals across denoising steps.
- It beats existing sparse-attention baselines on fidelity and speedup: 1.80x denoising speedup on Minimax-H3-Base and 2.32x on 3D asset generation vs dense attention, with negligible quality loss.
More from Multimodal
- Scanography-style portraits on Midjourney v8.2, full prompt included — tisch_eins · 2026-10-09
- Both influencers are AI: Higgsfield Katana makes synthetic creators trivial — CurieuxExplorer · 2026-10-09
- Claude Code Drives ComfyUI with Krea2 and LTX 2.5 in Hilarious AI Video Test — TheHollywoodGeek · 2026-10-09
- 20 GitHub repos turn Claude into a motion design studio — Roger_M_Taylor · 2026-10-09
- OmniCapBench: 786-video benchmark exposes weak long-horizon audio-visual reasoning — Tencent-Hunyuan · 2026-10-09
- Setting camera paths freely inside generated scenes is a surprisingly cool experience — XRarchitect · 2026-10-09