SGF+ decouples denoising and context-writing gradients, enabling 24-hour video from 5s training
Zihan Su · hf · 2026-10-08
SGF+ addresses systematic negative gradient alignment between denoising current frames and writing KV context in autoregressive video generation. It assigns separate parameters to the two roles while preserving interaction via causal attention, optimizing jointly with the original objective—no auxiliary losses. Trained on only 5-second rollouts, SGF+ supports continuous generation up to 24 hours without long-video fine-tuning, improving visual quality and long-horizon consistency over baselines in both framewise and chunkwise modes.
More from Multimodal
- AI-generated feature 'A Woman Asleep' enters major film festival's main competition — lmoroney · 2026-10-08
- vLLM-Omni technical report: a unified serving runtime for omni-modal generation — vllm_project · 2026-10-08
- Alaskan Raven Couple 'Conversing' Video Goes Viral on X — ZeroStateReflex · 2026-10-08
- Band Builds Audio-Reactive WebGL + Local SD 1.5 Pipeline for Live Improv Music Video — XploitXploit · 2026-10-08
- Claude turns OpenAI's 198-page Erdős conjecture proof into a 2-minute narrated 3D animation — imjustnewatai · 2026-10-08
- AI Short Film About Relationships Made With ComfyUI Agent Driver and Multi-Model Pipeline — TheHollywoodGeek · 2026-10-08