Running MiniMax H3 Fully Local on a 16GB GPU: 8-Step Video Generation & QA Lessons
Short_Regular_7191 · reddit · 2026-08-08
A developer shared a hands-on guide to running the 33B omni-modal MiniMax H3 video+audio model entirely locally on a 16GB RTX 5060 Ti.
Stack & Optimization
- Using ComfyUI's native support, NVFP4 pruning, and system RAM offload, the workflow bypasses VRAM limits smoothly.
- By utilizing a Turbo LoRA and a dual video/audio clock sampler, sampling steps were reduced from 20 to 8 steps. Generating a 5.2s native 1344x768 video takes just 11.4 minutes.
Automated Metrics vs. Human Eyes
- An automated QA script picks the best generation based on Whisper transcription accuracy, face similarity, and sharpness (variance of Laplacian).
- The script favored a clip with higher micro-contrast, but human review showed the lower-scoring clip was significantly more lifelike with natural motion. The new rule: whenever face similarity and sharpness metrics conflict, trigger manual human review.
Upscaling: Less is More
- Compared plain Lanczos upscaling against the SeedVR2 3B AI upscaler. For already high-res native generations, AI upscaling added artificial micro-contrast, making the result look "etched." A simple 1.4x Lanczos + crop won visually.
More from Infra
- Optimizing Small LLMs: Why the Standard Playbook Fails Below 1.5B Params — oli266 · 2026-08-08
- Keras Creator's AMD Bet Pays Off: Up 250% on Inference Compute Boom — fchollet · 2026-08-08
- Running Qwen3.6 27B on Tesla V100: 128K Context Config & Performance — Traditional_Bell8153 · 2026-08-08
- Running MiniMax H3 Locally on an RTX 3060: A Hands-on Test — the_frizzy1 · 2026-08-08
- Space-Based AI Data Centers? Startup Proposes 88,000-Satellite Constellation for 20GW Compute — VibeMarketer_ · 2026-08-08
- Inside the GB300 Rack: A Breakdown of Its 72-GPU Architecture — williamfalcon · 2026-08-08