MiniMax releases omni-modal generation model H3: 2K video at 1/3 the price of mainstream models
赛博禅心 · wechat · 2026-07-31
MiniMax has officially released H3, a general-purpose omni-modal generation model that understands and generates text, images, video, and sound, with native dual-channel audio and video output up to 15 seconds at 2K resolution. H3 excels in instruction following, text and brand presentation, and video-to-video motion transfer, making it suitable for advertising, e-commerce, gaming, and more.
Technical highlights:
- Contextual Omni Representation: Enhances multi-modal context understanding, unifying task descriptions via language.
- H3-VAE: High compression yields 4x sequence length efficiency, reducing training and inference costs.
- H3-OmniTransformer: Heterogeneous architecture for understanding and generation, boosting end-to-end training throughput by 30%.
- In-context Regeneration: Uses the base model to regenerate low-res outputs, surpassing traditional super-resolution.
Pricing: Under 1/3 the per-second cost of mainstream models at 2K, and under 1/2 at 768P.
Open-source plan: Weights to be released in the coming days, with compatibility for domestic chips.
Design philosophy: Breaking task boundaries, moving from specialized to general, unifying image, video, and audio generation tasks.
Related event: MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video(11 posts)→
More from Multimodal
- Testing Flux 3: Generating Synchronized Split-Screen Videos via Complex Prompts — umesh_ai · 2026-07-31
- MPIE-Bench: Evaluating Anatomical Errors in Multi-Person Image Editing — muset-ai · 2026-07-31
- RefCaptioner: Grounding Video Captions to Multiple Reference Images — KlingTeam · 2026-07-31
- Codex Launches Dedicated Image Agent UI: Supports Erasing, Resizing, and Batch Editing — op7418 · 2026-07-31
- Codex Launches Dedicated Image Agent UI with Batch Editing — op7418 · 2026-07-31
- MiniMax H3 Tops Video Editing Leaderboard, Undercutting Rivals at $7.80/min — ArtificialAnlys · 2026-07-31