CLSS: closed-loop streaming synthesis enables arbitrary-length LTX-2.3 video on a 16GB RTX 3080

nazgut · reddit · 2026-08-26

Video diffusion transformers only generate a few seconds per pass; naive timeline chunking collapses within a few hundred frames as exposure-bias drift compounds. Reddit user nazgut's CLSS (Closed-Loop Streaming Synthesis) treats the chunk hand-off as a feedback loop: chunks share a streaming latent buffer (SLB) overlap, keeping latent memory at O(overlap) instead of O(length), with lightweight inter-chunk corrections that fight drift without modifying any transformer weights.

Open-sourced as ComfyUI nodes: ComfyUI-LTX2.3-CLSS. The author demoed multi-scene prompt-following text-to-video on a single RTX 3080 (16GB VRAM) using the ltx-2.3-22b-dev-UD-Q4KS.gguf quant, 10 seconds per chunk; audio is still in progress.

Related event: CLSS: Closed-Loop Streaming Synthesis Enables Unlimited-Length Audio-Video Generation on a Single RTX 3080(2 posts)→

Original post →

More from Multimodal

Multimodal channel →