CtrlCache Speeds Up Interactive Video World Models 1.21–1.41x Without Retraining
Shangye Song · hf · 2026-10-07
- Interactive video world models generate chunk-by-chunk autoregressively, with each chunk needing several costly denoising steps. Existing training-free caching policies ignore control signals.
- Key insight: control inputs arrive before a chunk is denoised, so control-derived schedules cost zero extra forward passes; structural similarity drops around action changes while low-frequency structure persists.
- CtrlCache labels chunks as initial/transition/turning/steady, keeps full computation for the first two, and reuses transformer residuals for the latter two; a frequency-mixed history prior guidance adds information from the previous clean latent with no extra DiT forward.
- On Matrix-Game 2.0 and LingBot-World v1/v2, it achieves 1.21x–1.41x DiT-backbone speedups without retraining while improving WBench Overall on all three models.
More from Infra
- Musk denies slowdown: SpaceX accelerating AI data center buildout, exploring orbital compute — DimaZeniuk · 2026-10-07
- VEDA Sparse Attention cuts MiniMax H3 video gen time in half in ComfyUI with no visible quality loss — robomar_ai_art · 2026-10-07
- H200 vs multi-GPU RTX PRO 6000 Blackwell: how to pick inference hardware by budget — recentheartbroken · 2026-10-07
- Mistral release days: user reports speed slowed again with TPS around 30 — bdsqlsz · 2026-10-07
- EmbeddingGemma 2 ported to WebGPU: image-text photo search running fully in-browser — FinancialAd1961 · 2026-10-07
- Tracking LLM API model deprecations and rolling alias changes across providers — shamikhan005 · 2026-10-07