First/last-frame video workflow with GPT Image + LTX, author wants to move to local Qwen Image 2.1
ART-ficial-Ignorance · reddit · 2026-09-24
A creator shares an iterative abstract-video workflow: generating keyframes one by one in ChatGPT image generation's long context, then feeding first/last frames to LTX 2.3 for transitions. They propose a local pipeline with Qwen Image 2.1 — one style reference image plus 5 previous keyframes and an evolution prompt — and ask the community about Qwen's multi-reference consistency, continuity and the 'tendrils/spaghetti' degradation problem.
More from Multimodal
- Dreamina's rebuilt web app adds node-branching video workflow, $1.5 promo until Oct 9 — Div_pradeep · 2026-09-24
- DeepMind ships Gemini 3.8 Flash TTS with native multi-speaker overlap and laughter — kastnerkyle · 2026-09-24
- Treating AI video as a production pipeline: building a reusable cinematic producer agent in CREAO — FellMentKE · 2026-09-24
- Open-source ComfyUI node restores lost template filters for Local, Partner and Credit workflows — linus74RN · 2026-09-24
- Gemini 3.8 Flash TTS demo nails the Osaka auntie accent — heiga_zen · 2026-09-24
- ElevenLabs MCP adds voice, music, image and video generation to your AI assistant — aftahi_ai · 2026-09-24