MiniMax opens H3 on Hugging Face as a 33B multimodal video model with stereo audio
mark_k · x · 2026-08-04
MiniMax releases H3, an open-weights multimodal video model with stereo audio
MiniMax says H3 is now open on Hugging Face. It is a general-purpose omni-modal generative model that unifies text, images, video, and audio in one context.
- Generates video with native stereo audio and supports outputs up to 15 seconds at 2K.
- The company highlights strong instruction following, accurate text and brand rendering, and video-to-video motion transfer.
- It is aimed at advertising, branding, e-commerce, product design, and gaming.
- The post says H3 is a 33B-parameter Omni-Transformer conditioned on Qwen3-VL-32B.
- Checkpoints listed: FL2VA and Ref2VA; local inference works with Diffusers, SGLang, vLLM, and native ComfyUI support.
Related event: MiniMax Open-Sources 33B Multimodal Video Model H3(2 posts)→
More from Multimodal
- Dreamina launches Seedance 2.5 globally and claims the lowest price across platforms — manishkhosiya · 2026-08-04
- A ComfyUI workflow pushes Wan 2.2 image-to-video out to roughly 45 seconds — embryo10 · 2026-08-04
- A tiny H3 demo detail: the weights shift during the conversation — Oatilis · 2026-08-04
- Reddit users compare Qwen3.5-VL, InternVL and Gemma 4 for uncensored image captioning — TekeshiX · 2026-08-04
- AI-made “Ancient China” video draws attention for its stylized visual design — xiaosun86 · 2026-08-04
- Qwen Image Edit crashes in ComfyUI when blur nodes are added first — UnorthodoxyMedia · 2026-08-04