Running MiniMax-H3 Locally: Blind Comparison of Six Optimization Stacks
Primary-Confusion504 · reddit · 2026-08-11
To run the 33B-parameter, 385GB MiniMax-H3 video model smoothly on a single RTX 5090, the author compared six optimization stacks based on INT8 quantization and ConvRot.
- Scale: 28 prompts were used to generate 168 video clips under the same seed.
- Optimizations: Building upon SageAttention2 and FBCache, the comparison evaluates sparse attention (like Sol Attention), fewer sampling steps, and 4-8 step distillation LoRAs.
- Interactive Eval: The author built a web-based blind voting platform where users can pick the best quality with hidden labels or browse all 168 clips annotated with generation times.
More from Infra
- Optimizing NVFP4 Blockscaled GEMM on RTX Pro 6000 Blackwell — HanGuo97 · 2026-08-11
- CoreWeave Expands into APAC with 360MW Data Centers in Indonesia — Beth_Kindig · 2026-08-11
- OpenAI Believed to Cut Datadog Usage, Impacting Cloud Provider's Guidance — SumitGup · 2026-08-11
- Microsoft Plans to 'Significantly' Increase Production of Next-Gen AI Chips — thoefler · 2026-08-11
- fal Signs 3 Hot GenAI Model Companies, Expands H200 and B300 Capacity — gorkem · 2026-08-11
- Alphabet Aims to Raise $25B in Bonds to Fund AI Infrastructure Buildout — Beth_Kindig · 2026-08-11