ComfyUI-MiniMaxH3-CLSS update: per-scene refs, audio recompose, chunked latent upscale
nazgut · reddit · 2026-09-12
The author shipped a major ComfyUI-MiniMaxH3-CLSS update enabling single-go ref2vid with music via the MiniMax Music model. Key changes:
- Per-scene reference nodes: CLSSH3SceneReference(s) bind ref images/audio to one scene of a multi-scene generation (up to 9 images + 3 audios), with a 3-scene example workflow
- All-scene refs in one node: CLSSH3SceneReferencesAll attaches every ref to all scenes and auto-slices a soundtrack into per-scene windows (default 10s) — no rewiring needed
- globaltext on the prompt node: shared style/soundscape rules written once, prepended to every scene block
- Audio recompose: optional fresh audio takes from pure noise per chunk; measured that re-noising turbo audio regenerates the same take (cos 0.90–0.96), only fresh noise changes it
- Chunk-by-chunk latent upscale inside the sampler with cross-faded seams; long video never exists at high res at once (0.7 GB fp16)
- Fixes: recompose crash, bit-identical cross-chunk noise causality, audio telemetry bookkeeping
More from Multimodal
- Teaching image generation models to understand hex color codes, an ECCV 2026 paper — ducha_aiki · 2026-09-12
- Consistent multi-view 3D generation: using a multi-angle reference matrix — LCD-9 · 2026-09-12
- JD Cloud debuts Lingjing AIGC platform with infinite canvas; student made award-winning short film in 4 days — 量子位 · 2026-09-12
- Blender MCP with Astra delivers impressively smooth AI-driven animation — Dimillian · 2026-09-12
- Step-by-step Nano Banana workflow: split colorization into three passes to cut errors — aziz4ai · 2026-09-12
- Pwnisher winner tests full GPT Astra x Blender x Dreamina 3D animation pipeline — aziz4ai · 2026-09-12