FULL STORY

MiniMax H3: From Open-Source Release to Local Tests

MiniMax's open-source 2K video model H3 sparked massive community testing, proving capable of local execution on consumer hardware and impressing developers with its advanced multimodal capabilities.

2026-08-09 ~ 2026-08-13 · 4 episodes · 33 posts

Episode 1 · MiniMax Open-Sources H3 Video Model with 2K Native Generation (2026-08-09, 16 posts)

MiniMax has open-sourced its native 2K video generation model H3 (possibly codenamed H3 ref2va), which quickly sparked a wave of community tests due to its high generation quality and coherence. Current findings show the model can generate long videos, handle complex physical interactions, and run smoothly on local hardware, marking it as a strong contender in the multimodal video generation arena.

Confirmed

  • Core performance and open source: The model supports native 2K resolution generation with excellent detail and coherence. Developer ostrisai generated 1,000 videos covering a wide range of topics and released the 1.5-hour results as an open-source dataset.
  • Local deployment and long videos: User singularitynotnow confirmed local generation of a 30-second unedited video at 1024x576 resolution in 6 minutes 46 seconds, using 288GB VRAM with NVIDIA Sol-Attn acceleration. The Pixio team also confirmed breaking the 15-second limit with fast generation and good prompt coherence.
  • Multi-scenario creation: FDosha showcased realistic gravity manipulation; lazyspock used a long prompt to generate South Park-style 2D animation; Co-OB created a Dexter vs. Thanos animation; umeshai demonstrated macro creativity; Candiru666 showed cosmic color and light.
  • Workflow compatibility: jdude found H3 can seamlessly continue non-H3 generated clips, enabling mixed-model workflows. TheruleofThetra stitched three clips and used RIFE to interpolate to 96 FPS for a high-frame-rate effect.

Why it matters

  • On-device and open ecosystem: The high-quality open weights allow local running (e.g., Co-OB on RTX 3080), lowering the barrier for high-quality AI video creation and empowering independent creators and the open-source community.
  • Multimodal competitiveness: Breakthroughs in long video generation, complex prompt understanding, and multi-style rendering prove its top-tier capability in the competitive multimodal video track, offering new productivity tools for film, animation, and meme creation.

Episode 2 · MiniMax H3 Local Tests: 24-Min 2K Video on 16GB VRAM (2026-08-09, 12 posts)

Recent community tests of the MiniMax H3 video generation model confirm its local operation on consumer hardware and high controllability, with multiple advanced workflows emerging.

Confirmed

  • Local long-form video generation: Developers successfully ran the model locally on a single RTX 5060 Ti (16GB VRAM). Using NVFP4 quantization and lightx2v Turbo LoRA (8 steps) with Sol Attention, they generated a Brad Pitt talking video, a 24.4-minute native 2K (2048x1152) video, and a 2:47-minute micro-documentary on Socrates with 36 shots.
  • Advanced camera and scene control: Multiple users verified precise control. @HeftySide7892 used L2VA with structured prompts to recreate a coherent "moon landing set" and freeze on a reference image; @LanceCampeau shared an FFLF workflow for precise virtual camera movement via basic path prompts; @TinyTeam2511 and @Candiru666 reported stunning cinematic quality in commercial ads and complex close-ups.
  • Video editing workflow: @AltruisticTax1317 demonstrated local repainting and background replacement using SAM segmentation. Although audio-video coupling makes precise latent masking difficult, a developer has proposed a solution on GitHub.

Unconfirmed

  • Dynamic physics: @BestCandidate9060 noted in a 15-second test that static visuals are solid, but dynamic elements like fire are unnatural and shot transitions need improvement, indicating room for optimization in complex physical interactions.

Why it matters

These tests show high-end AI video generation moving to consumer local devices. Creators can run the model on 16GB VRAM, generate tens of minutes of 2K video, control virtual cameras, and perform local editing. This lowers the barrier for cinematic and commercial production, marking a shift from mere showcase to controllable industrial workflows.

Episode 3 · MiniMax Launches H3 Multimodal Video Model with ComfyUI Support (2026-08-10, 2 posts)

MiniMax launched H3, an open-source multimodal video model capable of processing text, image, video, and audio. Now integrated with ComfyUI, it supports various generation workflows including 15-second 768p video creation.

Episode 4 · MiniMax Video Models Impress Developers in Testing (2026-08-12, 3 posts)

MiniMax's latest video generation models have drawn high praise from developers for their stunning performance. Tests show these models excel in long video generation, complex camera movements, and scene reasoning, reportedly outperforming ByteDance's Seedance.