FULL STORY

MiniMax H3: From Open-Source Announcement to Top Benchmarks

MiniMax's omni-modal model H3 has progressed from open-source teasers to official release, topping video editing benchmarks. Early tests praise its native 2K audio-video generation, while significantly lowering deployment and generation costs.

2026-07-30 ~ 2026-08-03 · 13 episodes · 146 posts

Episode 1 · MiniMax H3 Video Model Tests Impress with Native 2K and Commercial Quality (2026-07-30, 59 posts)

MiniMax's latest video generation model H3 has entered early testing. Hands-on tests by multiple creators and developers show it reaches top-tier performance on several metrics, with capabilities considered on par with or even surpassing Seedance 2.0, demonstrating high commercial value.

Confirmed

  • Core specs: H3 generates videos up to 15 seconds, with native resolution up to 2560×1440 (2K). As an all-in-one model, it supports up to 12 cross-modal reference inputs (including images, audio, video, each up to 15 seconds).
  • Multimodal generation: Introduces a new Omni Reference workflow, breaking the traditional text-only limitation. Users can combine product images, reference videos, and audio for generation; even 5 static images can produce coherent short films or title sequences.
  • Generation capabilities: Multiple creators confirm H3's excellent complex prompt understanding, executing intricate camera movements (e.g., one-shot high-speed flythrough). It performs well in physical motion naturalness, action coherence, text rendering clarity, and high-fidelity speech with lip sync. It also incorporates color, lighting, and composition intent into generation logic rather than mechanically reproducing elements.

Why it matters

  • Commercial-grade quality: In side-by-side tests with identical prompts, multiple reviewers note H3's output surpasses Seedance 2.0 in product interaction dynamics, transition smoothness, and narrative pacing, approaching high-end commercial ads.
  • Workflow innovation: H3 can separate background and characters for replacement, and supports direct output without post-processing. Combining multimodal references, high-quality output, and low barrier to entry makes it a highly practical AI video tool.

39 more related posts →

Episode 2 · Google Rolls Out Free Gemini Video Generation and Editing (2026-07-30, 3 posts)

Google announced that users can now create up to 10 videos for free using Gemini until August 2026. The company also introduced Gemini Omni, a new tool offering advanced video editing features like object replacement and spatial interaction.

Episode 3 · MiniMax Releases Multimodal Model H3 with 2K Native Stereo Video (2026-07-30, 18 posts)

MiniMax officially released its multimodal generation model H3, which unifies understanding and generation of text, images, video, and audio, supporting native stereo audio output and up to 15-second 2K resolution (24fps) video. H3 is now available on Hailuo web, MiniMax API, and platforms like fal and Topview, and will soon be open-sourced with open weights. Its competitive pricing (under one-third of mainstream models) has drawn community attention.

Confirmed

  • Multimodal and output specs: H3 supports multimodal context understanding and can directly output videos with native stereo sound, up to 15 seconds at 2K resolution (24fps), with commercial-grade visual quality. Author @量子位 noted that the model breaks the limitation of traditional video models that only generate raw footage, integrating editing logic, typography, transitions, background music, and visual effects end-to-end to directly output 2K finished videos.
  • 12-Asset Reference: Allows users to combine up to 9 images, 3 video clips, and 3 audio clips as references, precisely locking character appearance, motion trajectories, and audio features.
  • Commercial-grade capabilities: Excels in precise text and brand rendering, video-to-video motion transfer (V2V Motion Transfer), etc., targeting commercial scenarios like advertising and e-commerce.
  • Platforms and pricing: The model is available on fal, Topview, and Hailuo. Author @赛博禅心 revealed that its API price is less than one-third of mainstream models; author @angrypenguinPNG also noted its cost is a fraction while matching Seedance's capabilities.

Unconfirmed

  • Open-source details and timing: Official and multiple creators confirm the model will be open-sourced soon, but the exact timeline is not fully set. Author @cocktailpeanut mentioned that officials said they would release weights "in compliance with laws and regulations" in the coming days; some netizens joked about putting it on BitTorrent due to impatience.

Why it matters

The release of H3 marks further maturity in multimodal fusion for video generation models. The combination of native stereo audio and high-quality 2K visuals, along with fine-grained multi-reference control, significantly lowers the barrier for commercial video production. Its competitive pricing and upcoming open-source release are expected to quickly capture market share in the AI video generation space.

Episode 4 · MiniMax H3 Model Coming Soon with Open Weights, Video and Multimodal Capabilities Spark Buzz (2026-07-30, 7 posts)

MiniMax officially revealed on Hugging Face that its next-generation H3 model is coming soon, with weights to be open-sourced shortly. The official account also showcased generation results on Hailuo AI, impressing the community. Early tests indicate the H3 video model supports native 2560×1440 (2K) resolution and up to 15-second video generation, fully model-generated without post-editing. Additionally, H3 emphasizes native multimodal understanding, allowing users to freely combine up to 12 reference materials (mixing video, text, image, and audio). Community discussions highlight that if H3 can run locally in environments like ComfyUI on an RTX 3060-class GPU, it would greatly benefit developers with limited compute resources. Furthermore, MiniMax's Hugging Face page also updated information on the M3 model, showcasing its capabilities in sparse attention and mathematical proof generation.

Confirmed

  • MiniMax officially stated on Hugging Face that the H3 model is coming soon and weights will be open-sourced quickly.
  • The official account showcased actual generation results based on the H3 model on the Hailuo AI platform, drawing community amazement.
  • Information about the MiniMax M3 model was also revealed, demonstrating its capabilities in sparse attention and mathematical proof generation.

Unconfirmed

  • Video model specs: According to early hands-on feedback, the H3 video model supports native 2560×1440 (2K) resolution and up to 15-second video generation, fully model-generated without post-editing.
  • Native multimodal capability: Some developers report that H3 focuses on native multimodal understanding, allowing users to freely combine up to 12 reference materials (supporting mixed video, text, image, and audio inputs).
  • Generation quality: Testers generated highly cinematic visuals with minimal prompts, and anime-style animation generation was impressive.

Why it matters

  • If H3 can run locally in environments like ComfyUI on an RTX 3060-class GPU, it would greatly benefit developers with limited compute resources, lowering the technical barrier.
  • Its strong multimodal mixed-input capability is seen as an important step toward benchmarking against industry frontiers in the multimodal domain.

Episode 5 · Gemini Omni Transforms Video Generation and Editing (2026-07-30, 3 posts)

Google's Flow studio integrated with Gemini Omni is revolutionizing video workflows via natural language. Users can now easily generate videos, change backgrounds, and auto-dub silent clips, showcasing the model's powerful creative potential.

Episode 6 · MiniMax H3 Tops Video Editing Chart, Announces Open Weights (2026-07-31, 7 posts)

According to the latest evaluation data from Artificial Analysis, MiniMax's latest video model H3 ties with Google's Gemini Omni Flash for first place in the audio-inclusive video editing category (Elo 1130 vs 1121), ranks second globally in text-to-video (with audio), and top three in image-to-video. MiniMax has officially confirmed plans to release the model weights under a community license, making it one of the few leading video models to announce open-sourcing.

Confirmed

  • In Artificial Analysis's video editing leaderboard (with audio), MiniMax-H3 (open-source) and Google's Gemini Omni Flash tie for first place by a narrow margin (Elo 1130 vs 1121)
  • Tied for second globally in text-to-video (with audio), top three in image-to-video
  • Model supports multimodal input, generating 5-15 second 24fps clips with native audio
  • Model supports native 2K resolution and stereo sound generation, with rich video editing features
  • 2K video priced at $7.8
  • MiniMax officially confirms plans to release model weights under a community license

Why it matters

  • H3 is one of the few leading video models to announce open-sourcing weights, potentially lowering the barrier to entry in video generation
  • Native video editing capability breaks the limitation of generation from scratch, supporting direct editing with text, images, etc.
  • The $7.8 price for 2K video is lower than most competitors, potentially attractive for commercial applications

Episode 7 · Topview Launches MiniMax H3 at 30% of Seedance Price (2026-07-31, 9 posts)

Video creation platform Topview announced a partnership with MiniMax (Hailuo AI) to launch the MiniMax H3 video generation model, becoming one of the first platforms to integrate it. The new model emphasizes native multimodal input and high-resolution output, with disruptive pricing aimed at providing high-volume creators with a more cost-effective option.

Confirmed

  • MiniMax H3 supports native 2K resolution output and can generate clips up to 15 seconds long.
  • The model understands text, images, audio, and video simultaneously, enabling cross-modal creative control and native audio-video output.
  • Pricing: H3 costs 70% less than Seedance 2.0, i.e., only 30% of its price.
  • Topview offers Ultra annual users 60 days of unlimited generation.

Why it matters

  • By integrating MiniMax H3, Topview gives creators a new option with solid resolution and duration at a significantly lower cost than competitors, potentially lowering the barrier and expense of high-quality video creation.

Episode 8 · MiniMax H3 Video Model Nears Release with ComfyUI Support (2026-08-01, 2 posts)

The upcoming MiniMax H3 open-source video model has been deeply optimized by ComfyUI to run on consumer hardware. Its workflow is already accessible in specific running nodes for users to test.

Episode 9 · MiniMax H3 Video Model Hands-on: 2K Quality and Impressive Multi-shot Consistency (2026-08-01, 12 posts)

MiniMax's upcoming open-source video generation model H3 has entered early testing, with several creators conducting hands-on tests via the Magnific platform. Feedback indicates outstanding performance in executing complex dynamic instructions, multi-shot visual consistency, and motion capture, along with extremely lenient censorship and a generation cost only a quarter of comparable models, positioning it as a strong competitor in video generation.

Confirmed

  • Core specs: H3 is an omni-model supporting up to 15-second video generation, native resolution of 2560×1440 (2K), native dual-channel audio, and simultaneous understanding of mixed image, video, and audio references.
  • Creator @Neggy5 noted that text-to-video quality is slightly blurry, but motion capture is very precise; image-to-video performs excellently at 2K, and the model has almost no content restrictions.
  • Creator @umeshai's tests show the model precisely understands and executes highly structured creative instructions, excelling in dynamic graphics, complex typography camera moves, and rhythm control. @卡尔的AI沃茨 also praised its effects and text stability.
  • Developer @gerardsans used H3 to create a game-style sequence animation, praising its full-reference capability to maintain high visual consistency across shots and fast transitions, resembling a real game trailer.
  • Creator @LudovicCreator showcased a demo turning a static poster into animation, calling it a great possibility for animation creation.
  • @mhdfaran tested complex long shots with elements like an ancient temple, glowing artifact, stone giant, and quick escape, achieving impressive multi-frame consistency with character, camera, lighting, water, and audio all fitting the same world.
  • The model is confirmed to run on Magnific, and a developer compared its lip-sync image-to-video effect with Grok.
  • In side-by-side tests, @Brojakhoeman and @evereveron78 compared H3 against LTX 2.3 and WAN using identical complex prompts like cyberpunk style.

Why it matters

  • MiniMax H3 demonstrates precise handling of complex creative instructions and advanced camera moves, especially breaking through in multi-shot consistency, directly addressing a core pain point in AI video generation. Its near-absence of censorship, low cost, and upcoming open-source release make it highly attractive to creators.

Episode 10 · ComfyUI Day-0 Support for MiniMax H3 Cuts VRAM by 66%, Runs on RTX 3060 (2026-08-01, 8 posts)

ComfyUI has officially announced Day 0 native support for the newly open-sourced MiniMax video model. Deep optimizations have drastically lowered the deployment barrier on consumer hardware, enabling smooth execution on standard personal computers.

Confirmed

  • Drastic Hardware Requirements Drop: The ComfyUI core team spent months optimizing the model, slashing VRAM usage by 66%. It now runs on a single GPU with 24GB of VRAM, and can even operate smoothly on an RTX 3060 paired with 32GB of system RAM.
  • Excellent Consumer-Grade Performance: Community tests show that using 8-bit weights on an RTX 3060, generating a 124-frame video at 832x480 resolution takes less than 10 minutes. The core team also successfully generated a 25-second 1080p video on a consumer-grade GPU.
  • Multimodal Generation Capabilities: The model supports text-to-video, image-to-video, and video editing. It can generate clips up to 2K resolution and 15 seconds long with stereo audio. The RunningHub version is now on Hugging Face, supporting image/audio/video-to-video+audio generation (e.g., T2VA, Ref2VA).
  • Comprehensive Ecosystem Support: A corresponding ComfyUI node library is now available for developers.

Why it matters

  • Breaking Hardware Barriers: The MiniMax video model is inherently massive. ComfyUI's native support and deep optimizations empower ordinary developers to experience and debug multimodal video generation locally without expensive, professional-grade GPUs. This significantly drives the adoption of open-source video models among individual creators.

Episode 11 · MiniMax Releases Open-Source Omni-modal Model H3 with Native 2K Audio-Video Generation (2026-08-02, 14 posts)

MiniMax has officially released the omni-modal generation system MiniMax-H3. The 33B parameter model features open weights on Hugging Face and is adapted for 🧨 Diffusers. It unifies multimodal context across text, image, video, and audio, supporting diverse inputs like text-to-video, image-to-video, and reference-image-to-video. Users can even combine video, image, text, and audio as reference materials in a single generation. For outputs, H3 natively generates end-to-end videos up to 2K resolution and 15 seconds long with native stereo sound, eliminating the need for post-dubbing.

Confirmed

  • Model Specs & Open Source: MiniMax-H3 has 33B parameters. Its weights are publicly available on Hugging Face, adapted for 🧨 Diffusers, with deployment solutions jointly launched by the vLLM project and MiniMax.
  • Generation Capabilities: Supports generating videos up to 2K resolution and 15 seconds long. It natively supports text/image-to-audio, achieving synchronized audio and video generation.
  • Leaderboard Performance: According to @NerdyRodent, H3 ranks first on the Artificial Analysis video editing leaderboard and second on the text-to-video leaderboard.

Why it matters

  • Audio-Visual Integration: H3 achieves coherent end-to-end generation of audio-visual content without post-dubbing, significantly lowering the barrier and complexity of video creation.
  • Open Source Ecosystem: As a rare and powerful open-source video model, H3 offers open weights and actively adapts to mainstream inference tools, providing the community with a highly competitive multimodal foundational solution.

Episode 12 · AI Video Model Generation Costs Compared (2026-08-02, 2 posts)

Creators tested multiple AI video models using identical prompts, revealing that MiniMax H3 generates 2K video with audio at a significantly lower cost per minute than Seedance models, reshaping production workflows.

Episode 13 · MiniMax H3 and FLUX3 Overcome Audio Hallucination (2026-08-02, 2 posts)

Recent tests show that new open-weight video models like MiniMax H3 and FLUX3 have resolved the persistent issue of audio hallucination, paving the way for advanced applications and bringing video generation closer to true world simulation.