Failed MiniMax H3 VAE detail experiment yields a useful 2X detail VAE and workflow
NoMouse9610 · reddit · 2026-10-01
A Reddit user released an attempt to improve the MiniMax H3 VAE, aiming to make 2X-upscaled pixels contain genuinely new detail rather than just an enlarged reconstruction.
The experiment
- Tried modifying the decoder, working around the patch/grid structure, detail residuals, and larger spatial decoders
- Tracing where fine information disappears showed the best detail signal exists before the H3 latent bottleneck
- Conclusion: a normally generated H3 latent simply doesn't contain that extra information — the original problem remains unsolved
What the failure produced (a 2-in-1 release)
- 2X VAE mode: works as a 2X video VAE for normal H3 generations
- 2X VAE + Detailed mode: for reference/image-to-video workflows, leveraging early encoder information from the source image before passing the enhanced reference into H3
The release includes the model, a ComfyUI node, workflows, and video comparisons, with the full research log — dead ends included — documented on Hugging Face.
More from Multimodal
- AI painting of the day, generated with ChatGPT Image 2.5 — DeryaTR_ · 2026-10-01
- Seedance 2.5 demo brings anime-level dual-sword choreography into photorealistic cinema — SimplyAnnisa · 2026-10-01
- Midjourney --sref 3896456162 recreates 1970s Kodak film & disco aesthetics — michaelrabone · 2026-10-01
- A finished LoRA run doesn't mean it learned the style: a reproducible SDXL validation workflow — no3us · 2026-10-01
- Editor open-sources open-fusion-mcp: Claude builds editable motion graphics inside DaVinci Resolve — JohnnyLegion · 2026-10-01
- AI-generated clip: Raven Noire writing in her diary while listening to goth music — Street-Pound5762 · 2026-10-01