Running MiniMax H3 video gen on 4GB VRAM: blurry hands at 512p or 6+ hours per clip at native res
unbenannt1 · reddit · 2026-09-08
A Reddit user documented days of experiments running MiniMax H3 Reference-to-Video locally on an RTX 3050 Laptop (4GB VRAM, WSL2, ComfyUI) to fix blurry hand rendering:
- Loading FL2V weights (minimaxh3fl2vaprunedint8) into the multi-reference node instead of the native Ref2V node fixed an earlier "ghost fingers" defect.
- At 512x288, hands still render as blurry blobs during gestures. Turbo distilled 4/8 steps unusable; off-label 10 steps made it worse; cfg=5.0 caused severe embossed cross-hatch artifacts; cfg=2.0 inconsistent; 30 steps no real gain.
- What works: 768x448 (H3's native short edge), no distillation, no CFG, native 20 steps — but 6h22m for one 15-second, 362-frame clip.
- Draft-then-upscale routes (pixel-space RealESRGAN, community latent 2x upscaler, tiled diffusion refine) each hit dealbreakers, including a ComfyUI core bug the author root-caused and patched locally.
Open questions: whether mid resolutions (608x352/640x384) capture most of the quality gain, whether turbo LoRAs are fundamentally tied to their distillation resolution, and whether any attention backend or quantization tricks make H3 realistic under 8GB VRAM.
More from Infra
- Jensen Huang confirms GPT-6 Astra trained on 100K+ Grace Blackwell NVL72 systems — rohanpaul_ai · 2026-09-08
- South Korea to give everyone free generative AI, backed by up to 512 B200 GPUs — IgorCarron · 2026-09-08
- Dev runs SDXL fine-tune fully on iPhone Neural Engine: 6-bit, 8 steps, offline — NovaDevCodeStudio · 2026-09-08
- Qwen 27B q8 vs bf16 on a DGX Spark: is the 1% token difference worth the memory? — superSmitty9999 · 2026-09-08
- Running Qwen3.8-27B for coding on 32GB VRAM — what local LLMs do you use and why? — theexile1337 · 2026-09-08
- A simple Windows tray app for monitoring NVIDIA GPU VRAM — Due-Committee-9591 · 2026-09-08