A video model trained on 5-second windows can now generate 4-minute clips
imjustnewatai · x · 2026-07-27
A video model trained on only 5-second windows can now generate up to 4 minutes of video.
- The method, called Self Gradient Forcing, lets future frames teach the model how earlier frames should store memory.
- It reportedly preserves the person, background, and layout much better than previous approaches over long generations.
- The post frames this as a step toward generated worlds that do not quickly drift, melt, or forget what happened.
Related event: Video Model Trained in 5 Seconds Generates 4-Minute Clips(2 posts)→
More from Multimodal
- Apertus 1.5 is a 70B European foundation model with native image and speech support — AxSaucedo · 2026-07-27
- Grok Imagine is being praised for natural motion and synced audio in AI video — XFreeze · 2026-07-27
- Designer builds beginner-friendly guides to FLUX, Stable Diffusion and ComfyUI — Masha-AI · 2026-07-27
- Local vs. Cloud: Evaluating image generation costs for indie game devs — Simple-Evidence-9125 · 2026-07-27
- Midjourney shows a new painterly style that turns people into blue-and-orange smoke — azed_ai · 2026-07-27
- Zombie Genesis trailer looks like a cinematic AI-generated apocalypse concept — HomeRemedyHealer · 2026-07-27