FULL STORY

MiniMax H3: From Early Testing to Open-Source Top

MiniMax's omni-modal model H3 progressed from early testing to official release and open-sourcing, topping video editing benchmarks and expanding to third-party platforms with its 2K native stereo video generation.

2026-07-30 ~ 2026-07-31 · 5 episodes · 84 posts

Episode 1 · MiniMax H3 Video Model Impresses in Comprehensive Tests (2026-07-30, 57 posts)

MiniMax (Hailuo AI) has launched early testing for its latest video generation model, H3, drawing massive attention for its outstanding capabilities and minimal restrictions. Tests show the model achieves top-tier industry performance, with many developers believing it rivals or even surpasses Seedance 2.0 and Kling Motion Control 3.0.

已确认

  • 核心规格: The H3 model generates videos up to 15 seconds long, with a native resolution of up to 2560 × 1440 (2K). As an all-in-one model, it supports up to 12 cross-modal reference inputs (including images, audio, and video, each up to 15 seconds).
  • 生成能力: Creators confirm H3 excels at understanding complex prompts and executing camera movements involving complex spatial transformations (like high-speed continuous tracking shots). It performs solidly in natural physical motion, action consistency, text rendering clarity, and high-fidelity lip-syncing. Furthermore, it integrates color, lighting, and composition intentions into its generation logic rather than mechanically reproducing elements.
  • 价格与生态: According to testers, H3 maintains low pricing and minimal usage limits while delivering high-resolution outputs. The companion Storyboard feature is also available.

为什么重要

  • 重塑商业片质感: Under identical prompts, reviewers noted H3's outputs surpass Seedance 2.0 in dynamic product interactions, transition fluidity, and overall narrative pacing, achieving a quality very close to high-end commercial ads.
  • 创作工作流革新: Using just a few static images (like 5 reference images or 5 static frames), H3 can generate coherent shorts or cinematic title sequences, perfectly isolating and replacing video backgrounds and characters. This combination of multimodal referencing, high-quality direct output, and a low barrier to entry makes it a highly practical AI video tool.

37 more related posts →

Episode 2 · MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video (2026-07-30, 15 posts)

MiniMax has officially released H3, an omni-modal generative model that unifies the understanding and generation of text, images, video, and audio. Its core breakthrough lies in native dual-channel audio output, capable of generating up to 15 seconds of 2K resolution video with commercial-grade visual quality and instruction-following capabilities.

Confirmed

  • Multimodal and Output Specs: H3 supports multimodal context understanding and can directly output videos featuring native stereo sound at up to 15 seconds long with 2K resolution.
  • 12-Asset Reference: Allows users to combine up to 9 images, 3 video clips, and 3 audio clips as references to precisely lock in character appearance, motion trajectories, and audio characteristics.
  • Commercial-Grade Capabilities: Performs exceptionally well in precise text and brand information rendering, as well as V2V Motion Transfer, meeting commercial viability standards.
  • Platform and Pricing: The model is now available on the fal platform. According to author @赛博禅心, its API price is less than a third of mainstream models.

Why It Matters

The release of H3 marks a further maturation of multimodal integration in video generation models. The combination of native stereo audio and high-quality 2K visuals, paired with granular multi-asset reference controls, significantly lowers the barrier to entry for commercial video production. Meanwhile, its highly competitive pricing strategy positions it to rapidly capture market share in the current AI video generation landscape.

Episode 3 · MiniMax H3 Model Announced with Open-Source Weights and Early Tests (2026-07-30, 7 posts)

MiniMax has officially announced the upcoming release of its next-generation model, H3, confirming that the model weights will be open-sourced soon. Early tests indicate exceptional performance in video generation and multimodal understanding, sparking intense interest and discussion within the community.

已确认

  • MiniMax revealed on Hugging Face that the H3 model is coming soon and its weights will be open-sourced shortly.
  • The official account showcased actual generation results using the H3 model on the Hailuo AI platform, stunning the community.
  • Information regarding the MiniMax M3 model was also disclosed, highlighting its capabilities in sparse attention and mathematical proof generation.

尚未确认

  • Video model specs: According to early feedback, the H3 video model supports a native resolution of 2560 × 1440 (2K) and can generate videos up to 15 seconds long, entirely raw output without post-editing.
  • Native multimodal capabilities: Developer tests suggest H3 focuses on native multimodal understanding, allowing users to freely combine up to 12 reference materials (mixing video, text, image, and audio).
  • Generation quality: Testers generated cinematic visuals using minimalist prompts, with stunning performances in anime-style animation generation.

为什么重要

  • If H3 can run locally in environments like ComfyUI using an RTX 3060 level GPU, it will greatly benefit developers with limited compute power and lower technical barriers.
  • Its robust multimodal mixed-input capability is regarded as a major step toward rivaling industry frontiers in the multimodal domain.

Episode 4 · MiniMax-H3 Open-Source Model Tops AI Video Editing Charts (2026-07-31, 3 posts)

MiniMax's newly open-sourced H3 video model tied with Google's Gemini for first place in Artificial Analysis's video editing rankings. The model supports multimodal inputs with native audio generation at highly competitive pricing.

Episode 5 · MiniMax H3 Video Model Lands on Topview at Competitive Pricing (2026-07-31, 2 posts)

The MiniMax H3 video generation model is now available on Topview, featuring native 2K resolution, 15-second generation, and cross-modal understanding. It is priced competitively at only 30% of the cost of Seedance 2.0.