MiniMax H3 opens its weights as a multimodal video model with native stereo audio
mark_k · x · 2026-08-04
MiniMax H3 opens its weights as a unified multimodal video model
MiniMax H3 is now open on Hugging Face, and the post describes it as a next-generation open-weights multimodal video model that combines text, images, video, and audio in one context.
Key details:
- Generates video with native stereo audio up to 15 seconds at 2K.
- Strong at instruction following, text rendering, brand rendering, and video-to-video motion transfer.
- Targeted at advertising, branding, e-commerce, product design, and gaming.
- The post claims 2K pricing is less than one-third of mainstream models.
- It uses a 33B-parameter Omni-Transformer conditioned on Qwen3-VL-32B.
- Available checkpoints include FL2VA and Ref2VA, with local inference support via Diffusers, SGLang, vLLM, and native ComfyUI.
The author calls it one of the strongest open video models released so far.
Related event: MiniMax Open-Sources 33B Multimodal Video Model H3(2 posts)→
More from Multimodal
- GitHub repo pairs Claude Code with Thrixel to build games from a single goal prompt — RanaHanocka · 2026-08-04
- ChatGPT Work turns a Backrooms prompt into a seamless 5-second looping video — goodside · 2026-08-04
- Ostris AI Toolkit now supports MiniMax H3 video-model training — RayHell666 · 2026-08-04
- MiniMax video model users ask how to match input resolution while praising its output quality — Careless-Constant-33 · 2026-08-04
- Minimax H3’s John Wick-style demo turns into a viral AI movie-moment — Parogarr · 2026-08-04
- A ComfyUI workflow pushes Wan 2.2 image-to-video out to roughly 45 seconds — embryo10 · 2026-08-04