MiniMax H3 Open-Weight Video Model Now in ComfyUI, Supports 15s 768p Clips
pmttyji · reddit · 2026-08-10
MiniMax H3 is an open-weight multimodal video generation model handling text, image, video, and audio. In ComfyUI, it enables text-to-video, image-to-video, first/last-frame, and reference-driven creation. H3 jointly generates visuals and synchronized stereo audio (dialogue, SFX, ambience, music) rather than adding audio later. Open weights support clips up to 15 seconds at 768p; hosted version supports up to 2K. The post discusses unified architecture, high-compression video representation for efficiency, and local setup tips.
More from Multimodal
- Zero cameras used: Creator generates entire 45-second short film with AI — heypearlai · 2026-08-10
- Treblo Launches AI Music Detector, Confirms First AI Song on Billboard Hot 100 — emmanuelvivier · 2026-08-10
- Fenix Flexin's Track Suspected of Using AI, Sparking Transparency Debate — emmanuelvivier · 2026-08-10
- Visual Comparison: H3 Int8 ConvRot vs W4A8_mixed Quantization — Devajyoti1231 · 2026-08-10
- Grok Imagine 2.0 Ships, Jumping to #2 Globally in Image Generation — eyishazyer · 2026-08-10
- SEEDANCE 2 Generates Sci-Fi Action Trailer for Zero-Gravity F1 Racers — tess-tipple · 2026-08-10