MiniMax-H3 Model Card Surfaces: Omni-modal Generation with Native Stereo Audio
ostrisai · x · 2026-08-03
The model card for MiniMax's new omni-modal generative system, MiniMax-H3, has appeared on Hugging Face. The system supports unified understanding of text, images, video, and audio, and can generate videos up to 15 seconds long in 2K resolution with native stereo audio.
Key specifications include:
- Output: 4-15 seconds duration, 24 FPS, 32 kHz stereo audio, supporting various aspect ratios from 21:9 to 9:16.
- Languages: Stable support for 11 languages including Chinese, English, Japanese, and Korean.
- Modes: Features multiple variants like first-and-last-frame generation (H3-Base-FL2VA), supporting text-to-video, image-to-video, and audio-to-video.
More from Models
- MiniMax-H3 Model Is Now Publicly Available — Sam_Witteveen · 2026-08-03
- China's LLM Race: Pricing, Licensing, and Feedback Loops to Decide the Winner — natolambert · 2026-08-03
- Small Models Show Surprising Potential, Giving Open Source Hope Against Closed Giants — max_paperclips · 2026-08-03
- Claude Spots 5-Year Hardware Wallet Bug in 8 Mins as AI Escapes Sandboxes — 新智元 · 2026-08-03
- Qwen3.8-Max Hits #4 on Frontend Code Arena, Ranking High Across All Domains — richie9830 · 2026-08-03
- Don't Trust Benchmarks Blindly: Expert Warns Harness Discrepancies Skew LLM Scores — cedric_chee · 2026-08-03