Local video generation can't hold 30-60s single-take clips: three continuation workflows all fail
SorryINeedHelp1 · reddit · 2026-09-08
A Redditor trying to generate a 30-60 second static single-speaker dialogue clip locally reports every continuation method degrades: latent continuation smudges colors and ruins skin texture past 30s; first-last-frame chaining shifts skin tone darker and dims backgrounds; ref2vid segmentation with a bridge clip shows the same color drift.
Even with high step counts and no turbo/attention shortcuts, 15 seconds is the OOM limit on decent settings. The poster asks whether a single uncut long take is simply not achievable with current local video models.
More from Multimodal
- GPT-6 Astra + Seedance 2.5 Pipeline Uses Blender as an Editable Middle Layer — rohanpaul_ai · 2026-09-08
- Try the GPT-6 Astra + Seedance 2.5 Blender-Pipeline Yourself on Dreamina — rohanpaul_ai · 2026-09-08
- "Sora Is Back": Viral Demo Shows Video Made with Astra — mallow610 · 2026-09-08
- Minimax H3 runs AI video locally at 720p — impressive but trails Seedance 2.5 — SimplyAnnisa · 2026-09-08
- fal extends 75% off H3 Max endpoints to Sept 15, adds real-time 1080p video — isidentical · 2026-09-08
- Dev turns every storefront he checked into at into photo-generated 3D miniatures — floguo · 2026-09-08