KAIST's GRACE cuts Wan2.1-I2V video generation latency by 11.1x with generation-aware latent compression

kaist-ai · hf · 2026-10-08

KAIST AI introduces GRACE, a two-stage generation-aware latent compression framework for speeding up video diffusion models. Naively compressing the autoencoder shifts latents away from what the pretrained DiT learned, requiring costly retraining or adaptation.

GRACE keeps a frozen base latent from the pretrained encoder while learning a residual latent for information lost under stronger compression, and aligns compressed latents with pretrained ones in the frozen DiT's feature space so the autoencoder optimizes for generation. The DiT is then adapted with lightweight fine-tuning and asymmetric denoising (base denoised before residual).

On Wan2.1-I2V-14B, GRACE cuts token count by 8x and latency by 11.1x at 480x832x81, while matching pre-compression quality on VBench.

Original post →

More from Infra

Infra channel →