Flux1 images to WAN2.2 clips: how to stitch them into 15+ second videos

wreck_of_u · reddit · 2026-08-24

A ComfyUI user describes their character video pipeline: a headless server batch-generates Flux1 images overnight using a reliable character LoRA, followed by manual filtering of body-horror and low-likeness results, then 5-7 second WAN2.1/2.2 clips via Replicate and Vast.

Their core question: can these short clips be concatenated into 15+ second videos with AI seamlessly extrapolating transitions? Should still images or short clips serve as input? Or would training a new LoRA from the images/videos to generate longer videos from text prompts be better? They also ask what tools to use now, mentioning Minimax H3.

Original post →

More from Multimodal

Multimodal channel →