MiniMax H3 Image-to-Video Workflow: Renders in 21 Mins on 12GB VRAM
vortis23 · reddit · 2026-08-22
Released a MiniMax H3 image-to-video ComfyUI workflow supporting 15-second multishot seamless stitching with full audio. By optimizing nodes and fixing audio bugs, render time on 12GB GPUs was reduced from 37 minutes to 21 minutes flat. The template is available on Civitai.
Related event: Running MiniMax H3 Locally on 12GB GPUs for Video Generation(2 posts)→
More from Multimodal
- Seeking Audio Upscaling LLMs: Is There a 'Super-Resolution' Model for Music? — LeatherRub7248 · 2026-08-24
- Describe your dream world to an AI dragon, which generates the planet for you — repligate · 2026-08-24
- Using kintsugi texture to fix cracks in edited 3D meshes — repligate · 2026-08-24
- Generating Hannibal Character Videos with FL2VA Model — Nimblecloud13 · 2026-08-24
- MiniMax H3 Revives Medieval Short Stories: Complete Workflow Shared — zanatas · 2026-08-24
- NAPE Audio Pretraining Achieves SOTA Without Decoders — kastnerkyle · 2026-08-24