SGF+ decouples denoising and context-writing gradients, enabling 24-hour video from 5s training

Zihan Su · hf · 2026-10-08

SGF+ addresses systematic negative gradient alignment between denoising current frames and writing KV context in autoregressive video generation. It assigns separate parameters to the two roles while preserving interaction via causal attention, optimizing jointly with the original objective—no auxiliary losses. Trained on only 5-second rollouts, SGF+ supports continuous generation up to 24 hours without long-video fine-tuning, improving visual quality and long-horizon consistency over baselines in both framewise and chunkwise modes.

Original post →

More from Multimodal

Multimodal channel →