ComfyUI MiniMax H3 tutorial: full T2V & I2V quick-start guide

Fun_Walk_4965 · reddit · 2026-08-27

A hands-on ComfyUI MiniMax H3 tutorial covering text-to-video and image-to-video quick start.

Files needed: minimaxh3fl2vaprunedint8convrot.safetensors for the diffusion model, a Qwen3-VL-32B NVFP4 AWQ text encoder, plus video and audio VAEs; an optional 8-step Turbo LoRA.

Key settings: start T2V at 1344×768, 24 FPS, 20 steps. The FL2VA checkpoint handles 0 images = T2V, 1 image = first- or last-frame I2V, 2 images = first+last frame generation. H3-Base uses a 768px short edge, 32px width/height alignment, and H3's temporal frame grid. H3 generates audio and video together—no sound usually means the audio VAE isn't wired to the final save node. Recommended order: T2V → single-image I2V → advanced features.

Original post →

More from Multimodal

Multimodal channel →