Kijai releases 4-bit quantized MiniMax H3, enabling low-VRAM video generation
CurrentNew1039 · reddit · 2026-08-08
Developer Kijai has released a w4a8 mixed-precision 4-bit quantized version of the experimental MiniMax H3 model. This version significantly reduces VRAM and RAM requirements, making it highly suitable for devices with 8GB of VRAM or less, while maintaining or even exceeding the inference speed of int 8.
The model is now open-sourced on Hugging Face, with core model files ranging from 11.8GB to 12.5GB. Additionally, it includes a highly compressed 2.9GB int 8 convrot video VAE, down from its 4.9GB fp16 counterpart. Users need to update ComfyUI to the latest version to run the model properly.
More from Infra
- Inflect Launches Free AI Crawler Index to Track LLM Web Scraping Activities — Scobleizer · 2026-08-08
- Why Are Consumer AI Devices Stuck? Barriers to On-Device LLM Chips — Waste-Intention-2806 · 2026-08-08
- NVIDIA Won't Sell High-VRAM Consumer GPUs; Inference Accelerators May Dominate — mike64_t · 2026-08-08
- Japanese Chemical Giant Monopolizes E-beam Source Material, Creating Key Chokepoint for China — PAstynome · 2026-08-08
- Cloudflare Open-Sources Computer: A Persistent Virtual Filesystem for AI Agents — bibryam · 2026-08-08
- Profiling LLM Inference with SGLang: Identifying Production Bottlenecks — BanghuaZ · 2026-08-08