KAIST's LongTake: long-horizon teacher forcing keeps 30-60s video generation dynamic

kaist-ai · hf · 2026-10-09

KAIST AI introduces LongTake, a two-stage pipeline for autoregressive video diffusion that sustains dynamics over long rollouts. Its Long-Horizon Teacher Forcing trains the model to predict later frames conditioned on long ground-truth video prefixes, extending supervision beyond the short training horizon and directly strengthening DMD initialization without an intermediate few-step distillation stage. Under the same 5-second DMD setup, it yields substantially higher dynamic degree than short-horizon TF at comparable aesthetics on 30s rollouts; Hybrid DMD attains the highest dynamic degree among evaluated methods at both 30s and 60s, sitting on the dynamics-aesthetics Pareto front.

Original post →

More from Multimodal

Multimodal channel →