First/last-frame video workflow with GPT Image + LTX, author wants to move to local Qwen Image 2.1

ART-ficial-Ignorance · reddit · 2026-09-24

A creator shares an iterative abstract-video workflow: generating keyframes one by one in ChatGPT image generation's long context, then feeding first/last frames to LTX 2.3 for transitions. They propose a local pipeline with Qwen Image 2.1 — one style reference image plus 5 previous keyframes and an evolution prompt — and ask the community about Qwen's multi-reference consistency, continuity and the 'tendrils/spaghetti' degradation problem.

Original post →

More from Multimodal

Multimodal channel →