ComfyUI video generation experiments: image and audio reference workflow

TensorTinkererTom · reddit · 2026-08-26

The author shares MM (reference image and audio) experiments created with ComfyUI. Despite rough editing and frequent audio glitches, the post details the workflow: primarily a 20-step out-of-the-box ComfyUI workflow with an audio node added. The Load Video node was used to grab the last 24 frames of the previous video to continue the sequence, with limited success. The author notes that a 4-step LoRA was introduced halfway through; even with low weight, audio glitches remained frequent, though video quality was generally decent. Links to sample clips and character swap prompts are provided via Pastebin.

Original post →

More from Multimodal

Multimodal channel →