MC-Sparse speeds MiniMax video denoising 1.80× at 15% attention density

MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers

Jiarui Chen, Zeqiang Lai, Jiangshan Wang, Ziheng Ouyang, Ye Huang, Xiangyu Yue, Cewu Lu, Chunchao Guo

cs.CV, cs.AI

2026-10-06

Training-free MC-Sparse caches exact KV picks and dense-sparse residuals. On MiniMax-H3, 15% attention density yields 1.80× faster denoising at 27.30 dB PSNR vs dense outputs.

What problem this solves

Video and 3D generation push Diffusion Transformer sequences into the tens of thousands of tokens. Attention grows with the square of that length, and every denoising step pays it again. Sparse attention keeps only a fraction of the query-key pairs, but near 30% density the outputs drift from dense attention.

Training-free systems usually implement this as block sparsity. Queries and keys are cut into GPU-tile blocks, scored with mean-pooled vectors, and kept or dropped whole. The approximate score and the block grid are tangled, so a quality drop cannot be charged cleanly to either one.

Matched-density oracles separate them. Each oracle keeps as much true attention mass as the budget allows. Fidelity rises when selection moves from whole blocks, to one shared set of KV tokens per query group, to a private set for every query. Against mean-pooled vanilla block-sparse attention, the remaining gap has three parts. Structural binding keeps irrelevant neighbors of an important token and forces dissimilar queries to share one selection. Selection error shows up because a softmax of the mean query is not the mean attention. The discarded tail still carries nonzero weight, and even a per-query oracle leaves that gap. All three appear on Wan2.1-1.3B at 480p.

Method

MC-Sparse targets those three errors and does not update weights. Grouping and exact selection run only on anchor steps. Later steps reuse the query groups, the KV indices, and an output residual.

KV tokens are chosen one by one and may cross block boundaries. The kernel gathers keys and values from indices instead of repacking them into blocks. Queries still have to fill a compute tile, so PDDP cuts similar queries into groups of exactly C, the tile width: project onto the principal direction, split at the median, and repeat. Equal groups avoid the padding that uneven k-means clusters add on top of the nominal budget.

For each group, the K selected keys are those with the largest summed attention weight. The anchor step finds that top-K with a two-pass kernel. FlashAttention does not expose logits, and the extra QK pass costs about half a dense attention GEMM. Reusing an earlier exact index set beats mean pooling recomputed every step. Q, K, and V still come from the current step.

Token-level cuts slice through blocks, so a block mean no longer describes what was dropped. The anchor step stores dense output minus sparse output. That residual changes less across denoising steps than the attention output itself. A reuse step adds the cached residual to a fresh sparse attention call. Groups, indices, and the residual refresh together. When GPU memory is tight, inactive cache entries move to CPU and the next layer is prefetched.

Results

Speedup covers DiT denoising only, warm-up included, condition encoding and VAE decoding excluded. PSNR, SSIM, and LPIPS measure agreement with the dense-attention output.

MiniMax-H3-Base, 768p text-to-video:

MethodDensityPSNR↑SSIM↑LPIPS↓Speedup
Sol-Attn32.4%23.660.8160.2341.59×
PISA30%23.450.8110.2431.52×
MC-Sparse25%28.440.8990.1611.61×
MC-Sparse15%27.300.8820.1761.80×

At 15% density MC-Sparse is sparser than Sol-Attn, about 3.6 dB higher in PSNR, and moves the speedup from 1.59× to 1.80×. Dense denoising takes 1062 seconds, against 660 seconds at 25% and 590 seconds at 15%. VBench ImgQual moves from 69.20 to 69.00 and 68.94; BgCons from 91.72 to 91.81 and 91.85.

HunyuanVideo-13B at 720p reaches 32.89 dB and 1.82× at 25% density, and 30.78 dB and 2.00× at 15%. Sol-Attn at 32.3% density sits at 28.06 dB and 1.79×. On Wan2.1-14B, MC-Sparse at about 25% density has the highest PSNR, the highest SSIM, and the lowest LPIPS among the sparse methods listed: 28.81 dB, 0.912, and 0.128 on text-to-video, and 32.11, 0.929, and 0.117 on image-to-video.

On the internal image-to-geometry model HY3D-Internal, 15% density gives a 2.32× denoising speedup and a Chamfer distance of 0.177 against the dense mesh. PISA at 25% density is at 0.976, and Sol-Attn at 36.1% density is at 1.901. Volumetric IoU is 82.91 and [email protected] is 96.33. Input-image consistency stays with the dense run: Uni3D-I 0.3311 and ULIP3D-I 0.1206, versus 0.3309 and 0.1204.

A cumulative ablation on Wan2.1-1.3B at 480p and density 0.2 starts from vanilla block-sparse attention at 20.96 dB. Reused exact selection reaches 21.60, token-level keys 22.57, query grouping 23.51, and residual compensation 27.05. The residual adds 3.54 dB in that stack, 1.25 dB on vanilla block-sparse attention, and 0.49 dB on PISA.

The attention kernel alone, excluding selection and grouping, is measured on Wan2.1-14B at 720p with 75,600 tokens. At density 0.2, MC-Sparse runs at 4.79× FA3, 96% of the ideal 1/s ratio, next to 99% for regular-block BSA, 82% for PISA, and 80% for SVG2. Spread over the full trajectory, anchor overhead equals about 5 to 6 extra points of density. Amortized Fast PDDP grouping costs about 1/3.96 and 1/6.14 of Flash-KMeans at 480p and 720p.

Why it matters

A trained video or 3D DiT can swap this cache into the denoising loop and leave the weights alone. Keeping 15% to 25% of the interactions lines up with roughly 1.6× to 2.3× faster denoising, while automatic perceptual scores stay close to the dense run.

The residual does not fix a weak selector by itself. PISA already compensates with block statistics, and stacking this residual adds under 1 dB. Token-level selection and equal-size groups have to clean up the error first. The tail correction then becomes the largest single gain.

Every anchor still runs full dense attention, and the amortized overhead is another 5 to 6 density points. On short sequences, where attention is not the main cost, the net win shrinks.

Limitations

The paper has no separate limitations section.

PSNR, SSIM, and LPIPS score similarity to dense outputs. At 15% density MiniMax is still at 27.30 dB, so a pixel gap remains. The two VBench scores barely move, and the paper reports no user study.

Densities in the main tables often do not match. MC-Sparse wins both PSNR and speed at a lower density than Sol-Attn and PISA, and matched-density head-to-heads are absent from those tables. On 3D, PISA's speedup is written as less than 1.87×, so it cannot be subtracted cleanly from 2.32×.

Component ablations stop at a 1.3B model and 480p. The large-model results do not break out each piece again. HY3D-Internal is not public, so the geometry numbers cannot be rerun outside.

Reuse assumes that attention and the residual change slowly on nearby steps. The plots extend the gap to 25 steps, and few-step distilled samplers are not tested. Residual-plus-index traffic is 2435 MB on Wan2.1-14B and 2710 MB on HunyuanVideo-13B. Prefetch hides the transfer behind compute, which still binds machines that are short on GPU memory or PCIe. Trainable efficient attention such as SANA, VSA, and SLA is outside the comparison.

Terms

Source

Related papers

All paper explainers