Testing Minimax H3: Multi-Reference Prompts Ignored, 5-Hour Generation on RTX 4090
RikkTheGaijin77 · reddit · 2026-08-05
A user testing the multi-reference model of Minimax H3 via ComfyUI encountered several issues. Firstly, the model completely ignored the provided face close-up reference, retaining the face from the full-body image instead. Secondly, it disregarded the fixed-camera setup of the reference dance video and autonomously added zooms and close-up shots.
Furthermore, the generation speed raised questions. Generating a 15-second, 1MP video took 5 hours and 43 minutes on an RTX 4090 with 128GB of RAM, heavily contrasting with community claims of generating similar videos in minutes on an RTX 4060.
More from Multimodal
- MiniMax H3 Reference-to-Video Test: Impressive Generation Quality — irmemon225 · 2026-08-05
- Custom ComfyUI Nodes Streamline MiniMax H3 Prompting and Media Loading — acedelgado · 2026-08-05
- Leaked WAN 3 Video Model Demo Tests Copyright Content Generation — ChrisGPT · 2026-08-05
- Running Full Flux Model on 12GB VRAM: A Local Inference Practice — TheRealFutaFutaTrump · 2026-08-05
- Anthropic Inks $10B Compute Deal; FLUX 3 Video Model Released — soumitrashukla9 · 2026-08-05
- SarasFlow: Open-Sourcing a Multilingual Educational Video Generator — sai_teja_ · 2026-08-05