MiniMax Details H3 Open-Source Model Architecture and Workflows in AMA

MiniMax 稀宇科技 · wechat · 2026-08-08

MiniMax's technical team hosted a Reddit AMA, providing in-depth answers regarding the core architecture, video generation limitations, and future plans of their open-source model H3.

Core Architecture & Training: H3 does not use MSA; its sparse attention strategy is closer to MoBA, currently applying 3D sparsification only to video tokens. To better control multimodal training balance, the model uses independent visual and audio VAEs. The team found that a model trained only on a simple next-frame prediction task showed strong zero-shot generalization in image editing.

Technical Limitations & Optimization: H3 is not natively trained for long trajectories of recursive continuations; it relies on reference conditioning (Ref2VA). The team is actively investigating issues like blurry textures and distorted distant faces in Ref2VA. Given the 60B parameter count, they are focusing on quantization, offloading, and step distillation to reduce inference costs.

Future Open-Sourcing: MiniMax plans to open-source a unified model for Text-to-Image and image editing sharing the same VAE encoder as H3. Additionally, the Regenerate-2K module will be open-sourced after efficiency and quality improvements.

Related event: MiniMax Hosts AMA for Open-Source Video Model H3(3 posts)→

Original post →

More from Multimodal

Multimodal channel →