SQuad reduces video generation attention compute by 67x with O(N√N) complexity

jm_alexia · x · 2026-08-19

The paper introduces SQuad, a framework achieving O(N√N) complexity for Self-Attention in video generation. Using a two-stage distillation process on the Wan 2.2 5B model, it matches the quadratic teacher's quality on VBench (83.20 vs 83.08) while cutting attention FLOPs by 67x and latency by 11x per step per block.

Original post →

More from Multimodal

Multimodal channel →