KAIST's LongTake: long-horizon teacher forcing keeps 30-60s video generation dynamic
kaist-ai · hf · 2026-10-09
KAIST AI introduces LongTake, a two-stage pipeline for autoregressive video diffusion that sustains dynamics over long rollouts. Its Long-Horizon Teacher Forcing trains the model to predict later frames conditioned on long ground-truth video prefixes, extending supervision beyond the short training horizon and directly strengthening DMD initialization without an intermediate few-step distillation stage. Under the same 5-second DMD setup, it yields substantially higher dynamic degree than short-horizon TF at comparable aesthetics on 30s rollouts; Hybrid DMD attains the highest dynamic degree among evaluated methods at both 30s and 60s, sitting on the dynamics-aesthetics Pareto front.
More from Multimodal
- Speridlabs Releases Iris-3B: Pixel-Space Generative Model That Can Replace DINOv2 — Apprehensive_Sky892 · 2026-10-09
- GPT-6 Luna Shows Surprising Skill at Spotting AI-Generated Images — Angaisb_ · 2026-10-09
- AI image of the day: GPT-6 Image 2.5 recreates 1940s Chinese village newsreel — DeryaTR_ · 2026-10-09
- Video character swap on 8GB VRAM: 640x480 render in 10:33 on RTX 4070 — big-boss_97 · 2026-10-09
- Google opens SynthID Detector globally; Vidu Q4 Preview prices video gen from $0.013/sec — 创业邦 · 2026-10-09
- Vigglorious Studio: chunked reference frames and guide keyframes fix character drift in long AI videos — Tablaski · 2026-10-09