Self Gradient Forcing adds missing memory-writing gradients for long video extrapolation

Junhao Zhuang · hf · 2026-07-23

Self Gradient Forcing adds missing gradient supervision for long video extrapolation

The paper identifies a historical context-gradient gap in autoregressive video diffusion trained with Self Forcing: future frames consume past key-value cache as frozen rollout state, so later losses cannot tell the model how earlier latents should be written into more useful memory.

What SGF does

Reported results

The authors say code and models will be released.

Original post →

More from Research

Research channel →