MiniMax H3 Open-Weight Video Model Now in ComfyUI, Supports 15s 768p Clips

pmttyji · reddit · 2026-08-10

MiniMax H3 is an open-weight multimodal video generation model handling text, image, video, and audio. In ComfyUI, it enables text-to-video, image-to-video, first/last-frame, and reference-driven creation. H3 jointly generates visuals and synchronized stereo audio (dialogue, SFX, ambience, music) rather than adding audio later. Open weights support clips up to 15 seconds at 768p; hosted version supports up to 2K. The post discusses unified architecture, high-compression video representation for efficiency, and local setup tips.

Original post →

More from Multimodal

Multimodal channel →