MiniMax Launches Omni-Modal Model H3 with Native Dual-Channel 2K Audio-Video Generation

智东西 · wechat · 2026-07-31

MiniMax has released MiniMaxH3, an omni-modal generative model capable of understanding text, images, video, and audio natively. It can generate dual-channel audio-video content up to 15 seconds long at 2K resolution.

The model breaks down traditional isolated pipelines (like text-to-image, image-to-video, voice cloning) by mixing multimodal data during pre-training for unified modeling. Users can freely combine different reference materials via natural language. MiniMax claims it ranks second on the ArtificalAnalysis text-to-video with audio leaderboard.

Pricing-wise, MiniMax states H3 costs less than 1/3 per second of mainstream models at 2K resolution. The company also plans to open-source the model weights in the coming days.

Related event: MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video(16 posts)→

Original post →

More from Models

Models channel →