MiniMax H3 Dropped: T2V, I2V, and Reference-to-Video with Audio

multimodalart · x · 2026-08-03

MiniMax has officially released the H3 video generation model. With 33 billion parameters, it supports text-to-video, image-to-video, and reference-to-video capabilities, complete with synchronized audio generation.

The model weights are now available on Hugging Face. It has been integrated into the 🧨 Diffusers and ComfyUI ecosystems, making it ready for deployment and experimentation on consumer-grade GPUs.

Related event: MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Stereo Video(13 posts)→

Original post →

More from Multimodal

Multimodal channel →