MiniMax H3 at max settings takes 26 min and 62GB VRAM per 15s clip on RTX Pro 6000
Realistic-Fennel-190 · reddit · 2026-09-30
A user benchmarked MiniMax H3 text-to-video at full settings on an RTX Pro 6000 (96GB): ComfyUI 0.38.0 native t2v template, 1280x736, 15s, 20 steps no turbo — 26.5 min per clip, 62,379MiB peak VRAM, 562W/81°C. Key details:
- Quantization: fl2va pruned int8 convrot + Qwen3VL 32B nvfp4 awq TE + video vae int8; only mxfp8 is emulated
- Model download: 39.5GB in 42 min
- A LoRA training attempt on H3 failed, so base model only
- Prompt was a 5-shot storyboard with timestamps; all shots appeared in order with audio generated in the same pass
- Turbo comparison: 8-step turbo runs 4.37 s/it (41s/clip), so max settings are 17x slower per step — author doubts 20 steps is worth it
- Template still demands the turbo LoRA on disk even with turbo off (2GB wasted)
More from Multimodal
- Dolphin AI launches agentic video studio with multi-shot character consistency, 45k test videos — thetripathi58 · 2026-09-30
- Dioramas open-sources a free 3D website framework with AI-generated assets and 20 example sites — Scobleizer · 2026-09-30
- Niantic Spatial demos a short film built inside a real-world Gaussian splat in hours — Scobleizer · 2026-09-30
- AI-reanimated Greta Garbo stars in SKF ball-bearing ad, panned as bland — nordicinst · 2026-09-30
- Dev predicts a universal DSL for video data will enable highly controllable world-simulation diffusion models — zeeshanp_ · 2026-09-30
- VideoLoop rewrites bounded working memory, hits 88.3% on VideoMME long video — Jinfa Huang · 2026-09-30