KAIST's GRACE cuts Wan2.1-I2V video generation latency by 11.1x with generation-aware latent compression
kaist-ai · hf · 2026-10-08
KAIST AI introduces GRACE, a two-stage generation-aware latent compression framework for speeding up video diffusion models. Naively compressing the autoencoder shifts latents away from what the pretrained DiT learned, requiring costly retraining or adaptation.
GRACE keeps a frozen base latent from the pretrained encoder while learning a residual latent for information lost under stronger compression, and aligns compressed latents with pretrained ones in the frozen DiT's feature space so the autoencoder optimizes for generation. The DiT is then adapted with lightweight fine-tuning and asymmetric denoising (base denoised before residual).
On Wan2.1-I2V-14B, GRACE cuts token count by 8x and latency by 11.1x at 480x832x81, while matching pre-compression quality on VBench.
More from Infra
- China's electricity glut turns data centers into a solution, as 14nm chips get pressed into service — teortaxesTex · 2026-10-08
- NAVER's DLoop Loops Speculative Decoding Before Verification, Gaining 5-41% Faster Inference Losslessly — naver-ai · 2026-10-08
- Transformer lead times balloon from 500 to 1,120 days, YC partner calls it a startup opportunity — ycombinator · 2026-10-08
- MIT's Christina Delimitrou uses AI to cut data center energy waste and downtime — nordicinst · 2026-10-08
- FT kicks off three-part series on China's breakneck AI infrastructure build-out, from Ulanqab to Shaoguan — zijing_wu · 2026-10-08
- Dev hits 1k tokens/sec prefill at 262K context with hybrid DeepSeek V4.1 Flash build — HankYeomans · 2026-10-08