Minima Quantizes 27B Hybrid LLM Fully to NVFP4 W4A4

Minima quantizes all 496 linear layers of Qwen3.8-27B to NVFP4 W4A4, shrinking the model 2.9x with near-lossless accuracy. The recurrent Gated DeltaNet components of hybrid LLMs prove surprisingly easy to quantize.

2026-09-04 ~ 2026-09-05 · 2 related posts