MiniMax H3 in ComfyUI: Generates 10s Video in 10 Mins on a 4090
wjc_5 · reddit · 2026-08-05
The author tested the newly open-sourced MiniMax H3 video model in ComfyUI, focusing on the full-reference workflow and local generation speed optimizations.
Performance Optimization
On a 48GB 4090, using the INT8 pruned model and NVFP4 text encoder to generate a 10-second video at 1 megapixel:
- Without acceleration: 30-34 minutes
- With SageAttention: 14-16 minutes
- With SageAttention + Prediction node: 10 minutes
Workflow & Tips
- The model supports up to 9 images, 3 videos, and 3 audio clips as reference inputs.
- The author noted that the Prediction node might cause visible softness on distant subjects. For high-motion scenes, it is recommended to use acceleration to find a good seed, then disable prediction for the final run.
- A system prompt template for web-based vision LLMs is provided to help organize rough stories and reference materials into precise H3-ready prompts.
Related event: ComfyUI Test: MiniMax H3 Generates Video on RTX 4090 in 10 Minutes(3 posts)→
More from Multimodal
- SenseTime Open-Sources 8B Multimodal Model SenseNova U1.5 — FellMentKE · 2026-08-05
- MiniMax Open-Sources Video Model: Community Runs It on 5GB VRAM in 48 Hours — DavidmComfort · 2026-08-05
- Decart AI Releases Anywear: Powering Mobile AR with Real-Time Video Models — bilawalsidhu · 2026-08-05
- xAI Video Models Hit Scenario: Supports Image-to-Video and Synced Audio — aziz4ai · 2026-08-05
- AI Video Meme: Generating a Black Hole Smith Eating Pizza — Moarkush · 2026-08-05
- Higgsfield Launches Seedance 2.5 with 7 Days of Free Unlimited Generations — SimplyAnnisa · 2026-08-05