Two-Stage Sampling Fixes MiniMax H3 Audio-Video Quality Trade-off
CornyShed · reddit · 2026-09-02
Addressing the common MiniMax H3 pain of balancing audio vs. visual quality, the author and u/LFAdvice7984 built a two-stage ComfyUI workflow: stage one generates audio, stage two the visuals, combined in output — each optimized independently with no trade-off. Supports text, first/last frame, or audio-to-video input (single-stage).
- Audio quality rates good to very good; default 50 steps excessive but reliability-first. Tip: select an output node and hit the play button to execute only up to it — audition audio before full generation.
- Visuals could improve with different turbo LoRA/sampler; large motion remains hard, with a possible third stage using different shift values.
- Generation times: 5s clip 10 min, 10s 25 min, 20s 65 min.
- Known issue: occasional prompt misunderstanding (ASMR, abstract clips) — unclear if prompt, enhancer, or model comprehension.
- The two-stage idea is model-agnostic, applicable to LTX 2.5 or upcoming Flux 3 Dev. Workflow JSON and prompts on Hugging Face; custom nodes include rgthree-comfy, KJNodes, RES4LYF, Spectrum-MiniMax-H3.
More from Multimodal
- WeMM-Embedding tops MMEB-v3: 9B scores 59.5, 2B beats every 7B/8B model — tomaarsen · 2026-09-02
- WeMM-Embedding-2B edges out Qwen3-VL-Embedding-8B on MMEB-v2 at quarter size — tomaarsen · 2026-09-02
- Tencent open-sources WeMM-Embedding: unified text/image/video embeddings, Apache 2.0 — tomaarsen · 2026-09-02
- Should style and speed LoRAs go into the latent upscale pass in two-pass workflows? — joseph_jojo_shabadoo · 2026-09-02
- MiniMax H3 at 768x causes melting faces in audio-synced singing videos — PersonalityLimp2593 · 2026-09-02
- LTX-2.5 under scrutiny: CTO Yaron Inger on whether open-weights video models really understand the world — kimmonismus · 2026-09-02