MiniMax Launches Omni-Modal Model H3 with Native Dual-Channel 2K Audio-Video Generation
智东西 · wechat · 2026-07-31
MiniMax has released MiniMaxH3, an omni-modal generative model capable of understanding text, images, video, and audio natively. It can generate dual-channel audio-video content up to 15 seconds long at 2K resolution.
The model breaks down traditional isolated pipelines (like text-to-image, image-to-video, voice cloning) by mixing multimodal data during pre-training for unified modeling. Users can freely combine different reference materials via natural language. MiniMax claims it ranks second on the ArtificalAnalysis text-to-video with audio leaderboard.
Pricing-wise, MiniMax states H3 costs less than 1/3 per second of mainstream models at 2K resolution. The company also plans to open-source the model weights in the coming days.
Related event: MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video(16 posts)→
More from Models
- DeepSeek-V4-Flash Repo Surfaces on Hugging Face with Million-Token Context — NielsRogge · 2026-07-31
- DeepSeek-V4-Flash Agent Eval: Completes 3D Task for $0.07 — cedric_chee · 2026-07-31
- DeepSeek-V4-Flash-0731 Model Weights Officially Released — shing3232 · 2026-07-31
- RL Training Could Unlock Massive Performance Gains for Kimi K3 and GLM 5.2 — airesearch12 · 2026-07-31
- Benchmarking Kimi K3, GLM 5.2, and DeepSeek V4 Pro in Agent Workflows — Teknium · 2026-07-31
- SAM 3.1 Quantized to INT4: 40% VRAM Savings with Identical Mask Quality — External_Quarter · 2026-07-31