Run MiniMax-H3 Locally: 4-bit Quantization Needs Only 8GB VRAM

petewoodbridge · x · 2026-08-05

ModelScope details how to run MiniMax-H3 locally. By using the 4-bit quantized version with DiffSynth-Studio, the model can run with a minimum of 8GB VRAM.

The solution also supports extreme mode on a 16GB Apple M-series Mac and LoRA training with a 24GB consumer GPU, all while preserving the model's original capabilities.

Related event: MiniMax-H3 Runs Locally on 8GB VRAM(3 posts)→

Original post →

More from Infra

Infra channel →