Minima Quantizes 27B Hybrid LLM Fully to NVFP4 W4A4
Minima quantizes all 496 linear layers of Qwen3.8-27B to NVFP4 W4A4, shrinking the model 2.9x with near-lossless accuracy. The recurrent Gated DeltaNet components of hybrid LLMs prove surprisingly easy to quantize.
2026-09-04 ~ 2026-09-05 · 2 related posts
- Why Gated DeltaNet survives 4-bit: full NVFP4 W4A4 quantization of a hybrid 27B LLM — minima-ai · 2026-09-04
- Minima quantizes all 496 layers of Qwen3.8-27B to NVFP4 W4A4, matching BF16 at 2.9x smaller — pbaylies · 2026-09-05