30B Video MoE Quantizations Tested: Most Users Should Wait
EntireBig7258 · reddit · 2026-08-03
The author compares the mainstream community quantization builds for the LingBot-Video 30B-A3B MoE model:
- GGUF route: Takes 17 GB on disk. By keeping weights in system RAM and dequantizing on demand, peak VRAM stays around 8 GB on an RTX 3070, though resolution and frame limits remain undocumented.
- SDNQ uint4 build: Exhibits severe quality degradation with a PSNR of 13.56 dB and visible color clipping.
- fp8 build: Only covers the 1.3B model, running about 1.27x faster than bf16.
- Diffusers native: FastVideo offers ports for both 1.3B and 30B (with refiner), while native ComfyUI support is currently a WIP.
Conclusion: The tooling is still stabilizing, and 30B quants have documented quality costs. Most users should stick to running the 1.3B model at full precision for now.
More from Infra
- Why SQLite's Architecture is Poised to Dominate the Agentic Era — glcst · 2026-08-03
- SK hynix Revenue Surpassing TSMC Raises Concerns Over Memory Boom Sustainability — jwt0625 · 2026-08-03
- DwarfStar runs DeepSeek V4 Flash on M3 Ultra at 37 t/s — mishig25 · 2026-08-03
- Cloudflare Kicks Off Agents Week: Building a Cloud Native for AI Agents — ritakozlov · 2026-08-03
- Open Video Models: License Restrictions and VRAM Requirements Compared — Mysterious_Sign_9501 · 2026-08-03
- AMD signs $14B+ deal with Core Scientific for 530MW AI data center capacity — Beth_Kindig · 2026-08-03