FULL STORY

MiniMax H3: From Open-Source Teaser to Top Benchmark

MiniMax released its fully open-source omni-modal model H3, delivering native 2K video with dual-channel audio. It quickly topped video editing benchmarks with high local deployability and low generation costs.

2026-07-30 ~ 2026-08-02 · 12 episodes · 127 posts

Episode 1 · MiniMax H3 Video Model Tests Impress with Native 2K and Commercial Quality (2026-07-30, 59 posts)

MiniMax's latest video generation model H3 has entered early testing. Hands-on tests by multiple creators and developers show it reaches top-tier performance on several metrics, with capabilities considered on par with or even surpassing Seedance 2.0, demonstrating high commercial value.

Confirmed

  • Core specs: H3 generates videos up to 15 seconds, with native resolution up to 2560×1440 (2K). As an all-in-one model, it supports up to 12 cross-modal reference inputs (including images, audio, video, each up to 15 seconds).
  • Multimodal generation: Introduces a new Omni Reference workflow, breaking the traditional text-only limitation. Users can combine product images, reference videos, and audio for generation; even 5 static images can produce coherent short films or title sequences.
  • Generation capabilities: Multiple creators confirm H3's excellent complex prompt understanding, executing intricate camera movements (e.g., one-shot high-speed flythrough). It performs well in physical motion naturalness, action coherence, text rendering clarity, and high-fidelity speech with lip sync. It also incorporates color, lighting, and composition intent into generation logic rather than mechanically reproducing elements.

Why it matters

  • Commercial-grade quality: In side-by-side tests with identical prompts, multiple reviewers note H3's output surpasses Seedance 2.0 in product interaction dynamics, transition smoothness, and narrative pacing, approaching high-end commercial ads.
  • Workflow innovation: H3 can separate background and characters for replacement, and supports direct output without post-processing. Combining multimodal references, high-quality output, and low barrier to entry makes it a highly practical AI video tool.

39 more related posts →

Episode 2 · Google Rolls Out Free Gemini Video Generation and Editing (2026-07-30, 3 posts)

Google announced that users can now create up to 10 videos for free using Gemini until August 2026. The company also introduced Gemini Omni, a new tool offering advanced video editing features like object replacement and spatial interaction.

Episode 3 · MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video (2026-07-30, 18 posts)

MiniMax officially released its multimodal generation model H3, which unifies understanding and generation of text, images, video, and audio, supporting native stereo audio output and up to 15-second 2K resolution (24fps) video. H3 is now available on Hailuo web, MiniMax API, and platforms like fal and Topview, and will soon be open-sourced with open weights. Its competitive pricing (under one-third of mainstream models) has drawn community attention.

Confirmed

  • Multimodal and output specs: H3 supports multimodal context understanding and can directly output videos with native stereo sound, up to 15 seconds at 2K resolution (24fps), with commercial-grade visual quality. Author @量子位 noted that the model breaks the limitation of traditional video models that only generate raw footage, integrating editing logic, typography, transitions, background music, and visual effects end-to-end to directly output 2K finished videos.
  • 12-Asset Reference: Allows users to combine up to 9 images, 3 video clips, and 3 audio clips as references, precisely locking character appearance, motion trajectories, and audio features.
  • Commercial-grade capabilities: Excels in precise text and brand rendering, video-to-video motion transfer (V2V Motion Transfer), etc., targeting commercial scenarios like advertising and e-commerce.
  • Platforms and pricing: The model is available on fal, Topview, and Hailuo. Author @赛博禅心 revealed that its API price is less than one-third of mainstream models; author @angrypenguinPNG also noted its cost is a fraction while matching Seedance's capabilities.

Unconfirmed

  • Open-source details and timing: Official and multiple creators confirm the model will be open-sourced soon, but the exact timeline is not fully set. Author @cocktailpeanut mentioned that officials said they would release weights "in compliance with laws and regulations" in the coming days; some netizens joked about putting it on BitTorrent due to impatience.

Why it matters

The release of H3 marks further maturity in multimodal fusion for video generation models. The combination of native stereo audio and high-quality 2K visuals, along with fine-grained multi-reference control, significantly lowers the barrier for commercial video production. Its competitive pricing and upcoming open-source release are expected to quickly capture market share in the AI video generation space.

Episode 4 · MiniMax H3 Model Coming Soon with Open Weights, Video and Multimodal Capabilities Spark Buzz (2026-07-30, 7 posts)

MiniMax officially revealed on Hugging Face that its next-generation H3 model is coming soon, with weights to be open-sourced shortly. The official account also showcased generation results on Hailuo AI, impressing the community. Early tests indicate the H3 video model supports native 2560×1440 (2K) resolution and up to 15-second video generation, fully model-generated without post-editing. Additionally, H3 emphasizes native multimodal understanding, allowing users to freely combine up to 12 reference materials (mixing video, text, image, and audio). Community discussions highlight that if H3 can run locally in environments like ComfyUI on an RTX 3060-class GPU, it would greatly benefit developers with limited compute resources. Furthermore, MiniMax's Hugging Face page also updated information on the M3 model, showcasing its capabilities in sparse attention and mathematical proof generation.

Confirmed

  • MiniMax officially stated on Hugging Face that the H3 model is coming soon and weights will be open-sourced quickly.
  • The official account showcased actual generation results based on the H3 model on the Hailuo AI platform, drawing community amazement.
  • Information about the MiniMax M3 model was also revealed, demonstrating its capabilities in sparse attention and mathematical proof generation.

Unconfirmed

  • Video model specs: According to early hands-on feedback, the H3 video model supports native 2560×1440 (2K) resolution and up to 15-second video generation, fully model-generated without post-editing.
  • Native multimodal capability: Some developers report that H3 focuses on native multimodal understanding, allowing users to freely combine up to 12 reference materials (supporting mixed video, text, image, and audio inputs).
  • Generation quality: Testers generated highly cinematic visuals with minimal prompts, and anime-style animation generation was impressive.

Why it matters

  • If H3 can run locally in environments like ComfyUI on an RTX 3060-class GPU, it would greatly benefit developers with limited compute resources, lowering the technical barrier.
  • Its strong multimodal mixed-input capability is seen as an important step toward benchmarking against industry frontiers in the multimodal domain.

Episode 5 · Gemini Omni Transforms Video Generation and Editing (2026-07-30, 3 posts)

Google's Flow studio integrated with Gemini Omni is revolutionizing video workflows via natural language. Users can now easily generate videos, change backgrounds, and auto-dub silent clips, showcasing the model's powerful creative potential.

Episode 6 · MiniMax H3 Tops Video Editing Chart, Announces Open Weights (2026-07-31, 7 posts)

According to the latest evaluation data from Artificial Analysis, MiniMax's latest video model H3 ties with Google's Gemini Omni Flash for first place in the audio-inclusive video editing category (Elo 1130 vs 1121), ranks second globally in text-to-video (with audio), and top three in image-to-video. MiniMax has officially confirmed plans to release the model weights under a community license, making it one of the few leading video models to announce open-sourcing.

Confirmed

  • In Artificial Analysis's video editing leaderboard (with audio), MiniMax-H3 (open-source) and Google's Gemini Omni Flash tie for first place by a narrow margin (Elo 1130 vs 1121)
  • Tied for second globally in text-to-video (with audio), top three in image-to-video
  • Model supports multimodal input, generating 5-15 second 24fps clips with native audio
  • Model supports native 2K resolution and stereo sound generation, with rich video editing features
  • 2K video priced at $7.8
  • MiniMax officially confirms plans to release model weights under a community license

Why it matters

  • H3 is one of the few leading video models to announce open-sourcing weights, potentially lowering the barrier to entry in video generation
  • Native video editing capability breaks the limitation of generation from scratch, supporting direct editing with text, images, etc.
  • The $7.8 price for 2K video is lower than most competitors, potentially attractive for commercial applications

Episode 7 · Topview Launches MiniMax H3 at 30% of Seedance Price (2026-07-31, 9 posts)

Video creation platform Topview announced a partnership with MiniMax (Hailuo AI) to launch the MiniMax H3 video generation model, becoming one of the first platforms to integrate it. The new model emphasizes native multimodal input and high-resolution output, with disruptive pricing aimed at providing high-volume creators with a more cost-effective option.

Confirmed

  • MiniMax H3 supports native 2K resolution output and can generate clips up to 15 seconds long.
  • The model understands text, images, audio, and video simultaneously, enabling cross-modal creative control and native audio-video output.
  • Pricing: H3 costs 70% less than Seedance 2.0, i.e., only 30% of its price.
  • Topview offers Ultra annual users 60 days of unlimited generation.

Why it matters

  • By integrating MiniMax H3, Topview gives creators a new option with solid resolution and duration at a significantly lower cost than competitors, potentially lowering the barrier and expense of high-quality video creation.

Episode 8 · MiniMax H3 Video Model Nears Release with ComfyUI Support (2026-08-01, 2 posts)

The upcoming MiniMax H3 open-source video model has been deeply optimized by ComfyUI to run on consumer hardware. Its workflow is already accessible in specific running nodes for users to test.

Episode 9 · MiniMax H3 Video Model Early Tests: Stunning 2K Output and Multimodal Consistency (2026-08-01, 12 posts)

The MiniMax H3 video generation model has entered early testing, with several creators getting a first look at its text-to-video and image-to-video capabilities. Initial feedback highlights strong performance in handling complex dynamic prompts, maintaining visual consistency across multiple shots, and capturing motion. The model also features relatively relaxed content moderation, positioning it as a highly competitive contender in the video generation space.

已确认

  • 要点 Creator @Neggy5 noted that while H3's text-to-video output is slightly blurry, its motion capture is highly accurate. The image-to-video performs exceptionally well at 2K resolution, and the model has virtually no content moderation restrictions.
  • 要点 Tests by creator @umeshai show the model accurately interprets and executes highly structured creative prompts, excelling in dynamic motion graphics, complex camera movements, and pacing control.
  • 要点 Developer @gerardsans used H3 to create a game-style cinematic sequence, praising its ability to maintain strong visual consistency across different shots and rapid transitions, delivering the look and feel of an authentic game trailer.
  • 要点 The model is confirmed to be available for testing via the Magnific platform, and developers have already conducted direct comparisons with Grok's image-to-video lip-sync capabilities.

为什么重要

  • 要点 MiniMax H3 demonstrates a precise ability to handle complex creative instructions and advanced camera work. Its breakthrough in multi-shot consistency directly addresses a core pain point in current AI video generation. Furthermore, its nearly restriction-free moderation and anticipated open-source release make it incredibly attractive to creators.

Episode 10 · MiniMax H3 Video Model Runs Locally on Consumer GPUs (2026-08-01, 3 posts)

MiniMax's newly released H3 video model demonstrates exceptional local deployment capabilities. ComfyUI developers successfully generated 1080p, 25-second videos in about 10 minutes on a consumer-grade RTX 3060 GPU, making high-quality video generation accessible to everyday users.

Episode 11 · AI Video Model Generation Costs Compared (2026-08-02, 2 posts)

Creators tested multiple AI video models using identical prompts, revealing that MiniMax H3 generates 2K video with audio at a significantly lower cost per minute than Seedance models, reshaping production workflows.

Episode 12 · MiniMax H3 and FLUX3 Overcome Audio Hallucination (2026-08-02, 2 posts)

Recent tests show that new open-weight video models like MiniMax H3 and FLUX3 have resolved the persistent issue of audio hallucination, paving the way for advanced applications and bringing video generation closer to true world simulation.