Local video generation can't hold 30-60s single-take clips: three continuation workflows all fail

SorryINeedHelp1 · reddit · 2026-09-08

A Redditor trying to generate a 30-60 second static single-speaker dialogue clip locally reports every continuation method degrades: latent continuation smudges colors and ruins skin texture past 30s; first-last-frame chaining shifts skin tone darker and dims backgrounds; ref2vid segmentation with a bridge clip shows the same color drift.

Even with high step counts and no turbo/attention shortcuts, 15 seconds is the OOM limit on decent settings. The poster asks whether a single uncut long take is simply not achievable with current local video models.

Original post →

More from Multimodal

Multimodal channel →