Running MiniMax H3 on a single RTX 5090: 0.8MP is the ceiling for 15s clips
Realistic-Fennel-190 · reddit · 2026-10-09
A Redditor tested MiniMax H3 image-to-video on a single 32GB RTX 5090 via ComfyUI's native template, holding duration at 15s and raising resolution until OOM:
- Setup: int8-quantized unet (20GB), qwen3vl 32b text encoder, int8 VAE, plus a lightx2v turbo LoRA at 8 steps; 90GB RAM cgroup limit supports dynamic weight offloading
- Results: 0.6MP in 476s, 0.7MP in 590s, 0.8MP in 746s; 0.9MP OOM'd at step 2/8 on an activation allocation even with dynamic loading
- Findings: VRAM sits near the top (31.7GB) at every resolution, so s/it is the better signal; prompt content doesn't change speed — only resolution and duration do
- Quality tradeoff: changing resolution changes the latent size and thus the noise, so each resolution is a new roll — 0.6 had the best motion/music but looked soft; 0.8 was best overall
- Prompt technique: timed phases (0-4s, 4-8s, 8-12s, 12-15s) with one small repeating move each
Untested: turbo off, --reserve-vram, other seeds. The author is asking how to hit 0.9MP at 15s on 32GB.
More from Infra
- Samsung expects 780% quarterly operating profit jump on AI boom — chemist_slime · 2026-10-09
- TRL v1.15 ships fused LM head: 82% less peak VRAM, 7x longer sequences — QGallouedec · 2026-10-09
- Dev streams a 66GB unquantized 22B video model on an iGPU with only 15.6GB shared RAM — Business_Swordfish_5 · 2026-10-09
- Asia-Pacific Needs $280 Billion to Build 26.5 GW of Planned Data Centers, Johor Leads — shashib · 2026-10-09
- Stepped MoE paper: one model scales from 1B to 4B parameters, beating dense counterparts by 2-5% — pmttyji · 2026-10-09
- DLoop: looped speculative decoding cuts target-model passes, boosting speedup 5-41% across EAGLE-3 and more — pmttyji · 2026-10-09